FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Site Reliability Engineer
ExpelPrincipal SRE securing Expel’s cloud-native cybersecurity platform. Leading Kubernetes, infrastructure, incident response, and reliability initiatives across distributed systems.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and maintaining platform features with a focus on reliability, networking, and cloud infrastructure. Proficient in infrastructure-as-code practices and experienced in mentoring teams while ensuring effective incident response and support for platform users.
Highest-signal resume keywords
Kubernetes OperationsGCP or AWS ExperienceInfrastructure-as-Code PracticesPython or Golang DevelopmentMonitoring and Observability Infrastructure
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Infrastructure-as-CodePythonGolangJavaScriptKubernetesCloud InfrastructureSystems OperationsIncident ResponseRoot Cause AnalysisMonitoring and Observability
Soft Skills
MentoringCollaborationCustomer-Minded ApproachMotivationTeamwork
Industry Keywords
Distributed EnvironmentsService ReliabilityPlatform SupportBlame-Free RetrospectivesSupport Rotation
Tech Stack
Tools & technologiesAWSCloudGoGoogle Cloud PlatformJavaScriptKubernetesLinuxPython
About the role
Key responsibilities & impact- Lead project work building and maintaining platform features across product reliability, networking, and cloud infrastructure
- Push infrastructure-as-code commits daily
- Occasionally write and test application code in Python, Golang, and JavaScript
- Mentor and motivate service owners on deploying, measuring, monitoring, and operating services at scale
- Participate in a weekly support rotation, including on-call pager duties and working-hours support for platform users
- Lead incident response, triage, and root cause analysis support
- Collaborate with architects and product stakeholders on reliability initiatives
- Pair-program with and mentor junior SREs
- Work from a shared backlog, pair-program weekly, peer-review work, and participate in blame-free retrospectives
Requirements
What you’ll need- Significant experience operating Kubernetes in highly distributed environments
- Experience running systems in GCP or AWS
- Exposure to monitoring and observability infrastructure and standard methodologies
- Understanding of infrastructure-as-code practices, tools, and patterns
- Some software development experience in Linux environments, preferably Python and/or Golang
- Six years of systems experience in operations or development
- Customer-minded approach supporting platform users and building organizational trust
- Collaborative disposition for working across teams
Benefits
Comp & perks- Bonus eligibility
- Equity
- Unlimited PTO
- Work location flexibility
- Up to 24 weeks of parental leave
- Excellent health benefits
- Reasonable accommodation for disabilities during hiring, work, and access to benefits