Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Resilient Co.

Senior Platform/Infrastructure Engineer

Resilient Co.

Senior Platform/Infrastructure Engineer at Resilient Co. deploying and maintaining Kubernetes and cloud-based infrastructure for platform reliability.

Posted 7/21/2026contractRemote • 🇦🇷 ArgentinaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in deploying and maintaining Kubernetes clusters, utilizing Python for automation, and implementing monitoring solutions with Prometheus. Proven ability to troubleshoot distributed systems and collaborate effectively across teams to enhance platform operations and support cloud-native migrations.

Highest-signal resume keywords
Kubernetes DeploymentPython AutomationPrometheus MonitoringCeph Storage SolutionsCloud Experience (AWS, Azure)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesPythonPrometheusCephAWSAzureBash ScriptingPostgreSQLOpenSearchFluent Bit
Soft Skills
CollaborationProblem-Solving
Industry Keywords
Site Reliability EngineeringPlatform EngineeringInfrastructure EngineeringDistributed SystemsCloud-Native Patterns

Tech Stack

Tools & technologies
AWSAzureCloudDistributed SystemsJavaKubernetesPostgresPrometheusPython

About the role

Key responsibilities & impact
  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain automation and tooling using Python to support platform operations.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

What you’ll need
  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.
  • Experience with OpenSearch (nice to have).
  • Proficiency with Bash scripting (nice to have).
  • Familiarity with Java-based services (nice to have).
  • Experience with Fluent Bit for log collection (nice to have).
  • Experience working with PostgreSQL (nice to have).

Benefits

Comp & perks
  • Laptop: BYOD.
  • Overtime Required: No.
  • Client Holidays (USA – Mandatory)