Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Lambda

Senior Site Reliability Engineer – Core Cloud Platform

Lambda

Senior Site Reliability Engineer at Lambda enhancing the reliability of AI cloud infrastructure across data centers. Focusing on Kubernetes operations, monitoring, and incident response within a collaborative environment.

Posted 7/29/2026full-timeSan Francisco • California • 🇺🇸 United StatesSenior💰 $240,000 - $356,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering, focusing on Kubernetes operations, infrastructure automation, and observability. Capable of leading incident response and improving system reliability through engineering solutions and effective communication across teams.

Highest-signal resume keywords
Site Reliability EngineeringKubernetes OperationsInfrastructure As CodeCI/CD WorkflowsObservability Platforms

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesTerraformGoPythonSLIsSLOsDistributed SystemsIncident ResponseAutomationMonitoring
Soft Skills
Clear CommunicationTeam CollaborationOwnershipSound JudgmentMentoring
Tools & Technologies
Argo CDFluxHelmKustomizeOpenTelemetryPrometheusGrafanaDatadog
Industry Keywords
InfrastructureDistributed SystemsProduction Software EngineeringPrivate CloudHybrid Cloud

Tech Stack

Tools & technologies
CloudDistributed SystemsFluxGoGrafanaKubernetesPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Operate and scale critical platform services across Lambda’s data centers.
  • Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.
  • Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.
  • Define SLIs, SLOs, error budgets, and operational readiness standards.
  • Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.
  • Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.
  • Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.
  • Lead production incident response, postmortems, and durable corrective actions.
  • Partner with Compute, Networking, Storage, Security, and Support teams.
  • Participate in on-call and improve its sustainability through automation and better tooling.
  • Mentor engineers and raise the reliability bar across the organization.

Requirements

What you’ll need
  • Have 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.
  • Have deep experience operating Kubernetes in production.
  • Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.
  • Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.
  • Are proficient with Terraform or similar infrastructure-as-code tools.
  • Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize.
  • Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.
  • Can build production-quality tooling in Go, Python, or a similar language.
  • Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure.
  • Have experience defining and operating against SLIs and SLOs.
  • Can lead effectively during high-severity incidents.
  • Approach recurring operational issues as engineering and automation problems.
  • Communicate clearly and work effectively across teams.
  • Bring strong ownership, sound judgment, and low ego.

Benefits

Comp & perks
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use