FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer, Cloud Platform
Salve.InnoSenior Site Reliability Engineer operating AWS and Kubernetes cloud platforms for mission-critical production services. Improving observability, automation, incident response, and reliability across engineering teams.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in maintaining production environments, enhancing observability solutions, and automating operational tasks. Proficient in incident response, troubleshooting, and promoting reliability engineering principles across teams.
Highest-signal resume keywords
Kubernetes OperationsAWS ExperiencePrometheus and GrafanaScripting in Bash, Python, or GoInfrastructure as Code with Terraform or Ansible
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesAWSPrometheusGrafanaBash ScriptingPython ScriptingGo ScriptingLinux AdministrationTerraformAnsible
Soft Skills
TroubleshootingCommunicationCollaborationProactive MindsetPassion for Automation
Tools & Technologies
ELK StackMySQLPostgreSQLRedisNoSQL Databases
Industry Keywords
Production ServicesContainer OrchestrationNetworking FundamentalsTCP/IPDNSLoad BalancingRoutingReliability EngineeringOperational ExcellenceContinuous Improvement
Tech Stack
Tools & technologiesAnsibleAWSDNSGoGrafanaKubernetesLinuxMySQLNoSQLPostgresPrometheusPythonRedisTCP/IPTerraformVoIP
About the role
Key responsibilities & impact- Maintain the reliability, availability, and performance of production and pre-production environments
- Monitor platform health and improve alerting, automation, and operational processes
- Respond to production incidents, participate in root cause analysis, and implement long-term improvements
- Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards
- Partner with software engineers to improve application reliability throughout the development lifecycle
- Develop and maintain operational documentation, troubleshooting guides, and runbooks
- Automate repetitive operational tasks to improve efficiency and reduce manual intervention
- Participate in on-call rotations while continuously improving incident response processes
- Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams
Requirements
What you’ll need- Bachelor's or Master's degree in Engineering, Computer Science, or a related field
- Strong experience operating Kubernetes or other container orchestration platforms
- Experience supporting large-scale production services
- Hands-on experience with AWS
- Experience with Prometheus, Grafana, and ELK
- Strong scripting skills in Bash, Python, or Go
- Experience administering Linux-based production environments
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible
- Solid understanding of networking fundamentals, including TCP/IP, DNS, load balancing, and routing
- Excellent troubleshooting, communication, and collaboration skills
- Proactive mindset with a passion for automation and reliability
- Experience with SIP or VoIP technologies (nice to have)
- Familiarity with MySQL or PostgreSQL (nice to have)
- Experience with Redis or other NoSQL databases (nice to have)
Benefits
Comp & perks- Long-term, full-time collaboration
- Flexible remote working environment
- Professional development opportunities, including training and technical learning
- Opportunity to work on innovative cloud technologies used by customers worldwide
- Collaborative engineering culture focused on knowledge sharing and continuous improvement
- Modern Apple equipment provided
- Inclusive, respectful workplace