FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer – Core Cloud Platform
LambdaSenior Site Reliability Engineer at Lambda enhancing the reliability of AI cloud infrastructure across data centers. Focusing on Kubernetes operations, monitoring, and incident response within a collaborative environment.
Posted 7/29/2026full-timeSan Francisco • California • 🇺🇸 United StatesSenior💰 $240,000 - $356,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering, focusing on Kubernetes operations, infrastructure automation, and observability. Capable of leading incident response and improving system reliability through engineering solutions and effective communication across teams.
Highest-signal resume keywords
Site Reliability EngineeringKubernetes OperationsInfrastructure As CodeCI/CD WorkflowsObservability Platforms
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesTerraformGoPythonSLIsSLOsDistributed SystemsIncident ResponseAutomationMonitoring
Soft Skills
Clear CommunicationTeam CollaborationOwnershipSound JudgmentMentoring
Tools & Technologies
Argo CDFluxHelmKustomizeOpenTelemetryPrometheusGrafanaDatadog
Industry Keywords
InfrastructureDistributed SystemsProduction Software EngineeringPrivate CloudHybrid Cloud
Tech Stack
Tools & technologiesCloudDistributed SystemsFluxGoGrafanaKubernetesPrometheusPythonTerraform
About the role
Key responsibilities & impact- Operate and scale critical platform services across Lambda’s data centers.
- Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.
- Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.
- Define SLIs, SLOs, error budgets, and operational readiness standards.
- Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.
- Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.
- Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.
- Lead production incident response, postmortems, and durable corrective actions.
- Partner with Compute, Networking, Storage, Security, and Support teams.
- Participate in on-call and improve its sustainability through automation and better tooling.
- Mentor engineers and raise the reliability bar across the organization.
Requirements
What you’ll need- Have 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.
- Have deep experience operating Kubernetes in production.
- Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.
- Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.
- Are proficient with Terraform or similar infrastructure-as-code tools.
- Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize.
- Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.
- Can build production-quality tooling in Go, Python, or a similar language.
- Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure.
- Have experience defining and operating against SLIs and SLOs.
- Can lead effectively during high-severity incidents.
- Approach recurring operational issues as engineering and automation problems.
- Communicate clearly and work effectively across teams.
- Bring strong ownership, sound judgment, and low ego.
Benefits
Comp & perks- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use