FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Platform/Infrastructure Engineer
Resilient Co.Senior Platform/Infrastructure Engineer at Resilient Co. deploying and maintaining Kubernetes and cloud-based infrastructure for platform reliability.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in deploying and maintaining Kubernetes clusters, utilizing Python for automation, and implementing monitoring solutions with Prometheus. Proven ability to troubleshoot distributed systems and collaborate effectively across teams to enhance platform operations and support cloud-native migrations.
Highest-signal resume keywords
Kubernetes DeploymentPython AutomationPrometheus MonitoringCeph Storage SolutionsCloud Experience (AWS, Azure)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesPythonPrometheusCephAWSAzureBash ScriptingPostgreSQLOpenSearchFluent Bit
Soft Skills
CollaborationProblem-Solving
Industry Keywords
Site Reliability EngineeringPlatform EngineeringInfrastructure EngineeringDistributed SystemsCloud-Native Patterns
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsJavaKubernetesPostgresPrometheusPython
About the role
Key responsibilities & impact- Design, deploy, and maintain production Kubernetes clusters and related services.
- Build and maintain automation and tooling using Python to support platform operations.
- Integrate and operate Prometheus for monitoring, alerting, and observability.
- Deploy and manage Ceph storage solutions for distributed workloads.
- Support platform modernization initiatives and migrate services to cloud-native patterns.
- Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
- Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
- Document platform designs, runbooks, and operational procedures.
- Participate in on-call rotations and incident response to maintain platform availability.
Requirements
What you’ll need- 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
- Proven experience deploying and operating Kubernetes in production.
- Strong Python skills for automation, tooling, and operational scripts.
- Experience implementing and operating Prometheus-based monitoring and alerting.
- Hands-on experience with Ceph or similar distributed storage systems.
- Cloud experience with AWS and Azure (designing, deploying, and operating services).
- Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
- Experience collaborating across teams to deliver platform improvements and migrations.
- Experience with OpenSearch (nice to have).
- Proficiency with Bash scripting (nice to have).
- Familiarity with Java-based services (nice to have).
- Experience with Fluent Bit for log collection (nice to have).
- Experience working with PostgreSQL (nice to have).
Benefits
Comp & perks- Laptop: BYOD.
- Overtime Required: No.
- Client Holidays (USA – Mandatory)