FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Customer Reliability Engineer - Infrastructure
AstronomerCustomer Reliability Engineer focusing on the reliability of Astronomer’s cloud infrastructure. Engaging with diverse customers to enhance operational reliability and product experience.
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Kubernetescloud infrastructureAWSGCPAzureLinuxObservability toolsPython scriptingDevOpsCI/CD
Soft Skills
troubleshootingcommunicationcustomer experienceproblem-solvingcollaborationprioritizationdocumentation enhancementcustomer engagementactive triagingguidance
Tools & Technologies
monitoring systemsalerting systemsdistributed systems
Industry Keywords
cloud-nativeproduction systemnetwork experiencecustomer solutionsoperational efficiencycustomer needspain pointsSLAsfully distributed teamcustomer documentation
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformKubernetesLinuxPython
About the role
Key responsibilities & impact- Provide solutions to customers to make them successful using our products.
- Troubleshoot Customer environments and engage in active triaging with customers
- Provide feedback to the product development teams on customer needs and pain points.
- Build out our monitoring and alerting systems.
- Build and maintain automation to ensure daily operational tasks are handled as efficiently as possible.
- Help direct the architecture of the products and contribute where possible.
- Own the customer experience, working directly with customers to prioritize and solve issues, meet SLAs, and provide “white glove” guidance on the path to production.
- Participate remotely within a fully distributed team.
- Enhance and Enrich customer documentation
- Work on a modern, sophisticated, cloud-native product that customers use to connect to dozens of other systems.
Requirements
What you’ll need- 5+ years of experience, preferably with large, complex cloud infrastructures operating at scale
- 3+ years of experience with Kubernetes
- Experience managing a Production distributed system with at least one major cloud provider (one or all: AWS, GCP, Azure)
- Strong Network Experience with one of the major Clouds
- Strong Linux experience
- Knowledge of how to operate and monitor issues for distributed systems
- Experience with Observability tools
- Previous experience in handling customers issues (internal and external)
- Strong Communication Skills
- DevOps or CI/CD experience
- Python scripting
- Good troubleshooting Skills
Benefits
Comp & perks- Help maintain 24x7 coverage through a specified 6-hour pager period during your work day.
- Participate in paid on-call rotation for weekend coverage.