Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Red Hat

Customer Site Reliability Engineer – OpenShift Managed Cloud Services, Kubernetes/AWS/Azure, Linux

Red Hat

Customer Site Reliability Engineer for Red Hat's OpenShift Managed Cloud, focusing on reliability and performance. Manage distributed systems and drive continuous improvement for customer satisfaction.

Posted 7/20/2026full-timeRemote • 🇮🇳 IndiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates advanced experience in managing large-scale distributed systems with a focus on OpenShift/Kubernetes, Linux-based systems, and enterprise configuration management tools like Ansible and Terraform. Proven ability to enhance system resilience, lead incident response, and mentor team members while maintaining superior communication with customers.

Highest-signal resume keywords
OpenShift/Kubernetes AdministrationLinux System ManagementEnterprise Configuration ManagementSoftware Engineering with GolangIncident Response Leadership

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
OpenShiftKubernetesLinuxAnsibleTerraformGolangPrometheusTCP/IP NetworkingCloud Platforms (AWS, Azure, GCP)Container Technologies
Soft Skills
Superior Communication SkillsMentoringProactive MindsetCollaborationCustomer-Focused Approach
Industry Keywords
Distributed SystemsSystem ResilienceIncident ResponseContinuous ImprovementCustomer Escalations

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudDistributed SystemsGoGoogle Cloud PlatformKubernetesLinuxOpenShiftPrometheusTCP/IPTerraform

About the role

Key responsibilities & impact
  • Manage large-scale, distributed systems, focusing on minimizing downtime and improving system resilience.
  • Maintain customer trust and confidence by ensuring stability and functionality of services.
  • Drive continuous enhancement of processes, tools, and methodologies to support the evolving needs of the service.
  • Lead the development of code and automation scripts to optimize the scalability, reliability, and performance of services.
  • Lead and participate in high-priority customer escalations, adopting a customer-first mindset.
  • Coordinate and execute complex incident response procedures, ensuring timely resolution and thorough postmortems.
  • Collaborate with cross-functional teams to enhance system robustness.
  • Demonstrate a proactive mindset to help preempt escalations and ensure reliable operations.
  • Document resolutions, root causes, and best practices to enrich the knowledge base and promote self-service solutions.
  • Mentor and coach team members, fostering a culture of continuous learning, knowledge sharing and collaboration.
  • Participate in on-call rotation and provide leadership during critical incidents.
  • Collaborate on strategic AI and automation projects designed to increase the efficiency of fleet operations and troubleshooting, ultimately delivering a better product experience for customers.

Requirements

What you’ll need
  • Advanced Experience with OpenShift/Kubernetes container platform support or administration.
  • Proficient with container-based technologies on Linux.
  • Proficient in managing Linux-based systems in a public cloud such as AWS, Azure, or GCP.
  • Advanced experience with enterprise systems monitoring; knowledge of Prometheus is preferred.
  • Advanced with enterprise configuration management such as Ansible, Terraform.
  • Software engineering experience using object-oriented languages; golang is preferred.
  • Superior communications skills and experience working directly with and presenting to customers.
  • Ability to quickly learn new technologies and follow industry trends.
  • Demonstrated ability to quickly and accurately troubleshoot systems issues.
  • Solid understanding of standard TCP/IP networking and common protocols.
  • Fluent in English and any additional language like Japanese, Chinese, Korean, Spanish is an advantage.

Benefits

Comp & perks
  • Health insurance
  • 401(k) matching
  • Flexible work hours
  • Paid time off
  • Professional development opportunities