Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
MyFitnessPal

Site Reliability Engineer

MyFitnessPal

Site Reliability Engineer improving MyFitnessPal’s production systems reliability and security. Engaging in incident response, observability, and infrastructure management.

Posted 7/26/2026full-timeRemote • 🇺🇸 United StatesMid-LevelSenior💰 $120,000 - $165,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering, focusing on SLI/SLO frameworks, incident response, and security integration within CI/CD pipelines. Proficient in managing Kubernetes and cloud infrastructure, with a strong emphasis on observability and policy-as-code practices.

Highest-signal resume keywords
Site Reliability EngineeringInfrastructure as Code (Terraform)Kubernetes ManagementCI/CD Pipeline Security IntegrationObservability Tooling (Datadog)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Go ProgrammingPython ProgrammingTypescript ProgrammingCloud Security FundamentalsPolicy-as-Code (Kyverno, OPA/Rego, Conftest)Incident Response LeadershipSLO-Driven Reliability PracticesVulnerability Triage and RemediationCapacity PlanningCloud-Cost Optimization
Soft Skills
JudgmentCommunication SkillsCoaching
Tools & Technologies
DatadogGitHub ActionsSASTDASTSCA
Industry Keywords
SOC 2PCI DSSHIPAAChaos EngineeringB2C Backend Environments

Tech Stack

Tools & technologies
AWSCloudGoKubernetesPythonTerraformTypeScript

About the role

Key responsibilities & impact
  • Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
  • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
  • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
  • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
  • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
  • Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
  • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
  • Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
  • Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
  • Coach team members and engineers across the org on reliability patterns and operational best practices

Requirements

What you’ll need
  • 5+ years in site reliability, platform, or infrastructure engineering
  • Strong programming skills for automation and tooling (Go, Python, Typescript or similar)
  • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
  • Proven track record leading incident response and building SLO-driven reliability practices
  • Working fluency with observability tooling (Datadog is a plus)
  • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
  • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
  • The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
  • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
  • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
  • Chaos engineering or game-day experience is a plus
  • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

Benefits

Comp & perks
  • healthcare
  • parental planning
  • mental health benefits
  • annual performance bonus
  • a 401(k) plan and match
  • responsible time off
  • monthly wellness and technology allowances