Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Expel

Principal Site Reliability Engineer

Expel

Principal SRE securing Expel’s cloud-native cybersecurity platform. Leading Kubernetes, infrastructure, incident response, and reliability initiatives across distributed systems.

Posted 8/14/2026full-timeRemote • 🇺🇸 United StatesLead💰 $167,300 - $242,600 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and maintaining platform features with a focus on reliability, networking, and cloud infrastructure. Proficient in infrastructure-as-code practices and experienced in mentoring teams while ensuring effective incident response and support for platform users.

Highest-signal resume keywords
Kubernetes OperationsGCP or AWS ExperienceInfrastructure-as-Code PracticesPython or Golang DevelopmentMonitoring and Observability Infrastructure

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure-as-CodePythonGolangJavaScriptKubernetesCloud InfrastructureSystems OperationsIncident ResponseRoot Cause AnalysisMonitoring and Observability
Soft Skills
MentoringCollaborationCustomer-Minded ApproachMotivationTeamwork
Industry Keywords
Distributed EnvironmentsService ReliabilityPlatform SupportBlame-Free RetrospectivesSupport Rotation

Tech Stack

Tools & technologies
AWSCloudGoGoogle Cloud PlatformJavaScriptKubernetesLinuxPython

About the role

Key responsibilities & impact
  • Lead project work building and maintaining platform features across product reliability, networking, and cloud infrastructure
  • Push infrastructure-as-code commits daily
  • Occasionally write and test application code in Python, Golang, and JavaScript
  • Mentor and motivate service owners on deploying, measuring, monitoring, and operating services at scale
  • Participate in a weekly support rotation, including on-call pager duties and working-hours support for platform users
  • Lead incident response, triage, and root cause analysis support
  • Collaborate with architects and product stakeholders on reliability initiatives
  • Pair-program with and mentor junior SREs
  • Work from a shared backlog, pair-program weekly, peer-review work, and participate in blame-free retrospectives

Requirements

What you’ll need
  • Significant experience operating Kubernetes in highly distributed environments
  • Experience running systems in GCP or AWS
  • Exposure to monitoring and observability infrastructure and standard methodologies
  • Understanding of infrastructure-as-code practices, tools, and patterns
  • Some software development experience in Linux environments, preferably Python and/or Golang
  • Six years of systems experience in operations or development
  • Customer-minded approach supporting platform users and building organizational trust
  • Collaborative disposition for working across teams

Benefits

Comp & perks
  • Bonus eligibility
  • Equity
  • Unlimited PTO
  • Work location flexibility
  • Up to 24 weeks of parental leave
  • Excellent health benefits
  • Reasonable accommodation for disabilities during hiring, work, and access to benefits