Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Grafana Labs

Senior Software Engineer – Databases, SRE

Grafana Labs

Senior Software Engineer focusing on customer reliability for Grafana Cloud databases. Designing solutions for SLOs and incident response, ensuring high reliability for complex environments.

Posted 7/20/2026full-timeRemote • 🇺🇸 United StatesSenior💰 $154,445 - $185,334 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on designing and implementing Service Level Objectives (SLOs) and automation for production reliability. Proficient in Kubernetes, cloud environments, and incident response processes, ensuring high availability and performance in complex customer systems.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Kubernetes ExperienceService Level Objectives (SLOs) DesignIncident Response ParticipationAutomation Implementation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesService Level Objectives (SLOs)GoPythonJavaLinux InternalsInfrastructure-as-CodeTerraformHelmJsonnet
Soft Skills
Problem-SolvingTroubleshootingIntellectual CuriosityTransparencyKindness
Tools & Technologies
AWSGCPAzureCloud StorageNetworking
Industry Keywords
Production EngineeringMulti-Tenant SystemsPost Incident Reviews (PIRs)Customer Reliability Engineering

Tech Stack

Tools & technologies
AWSAzureCloudGoGoogle Cloud PlatformJavaKubernetesLinuxPythonTerraform

About the role

Key responsibilities & impact
  • Partner closely with product engineering squads (embedded model)
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale our reliability practices
  • Ensuring our customers meet our SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serving as a primary escalation point and on-call for relevant incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews
  • Influence feature design to ensure production scalability and operability
  • Build automation to eliminate toil where needed
  • Improve alert quality and reduce noisy escalations

Requirements

What you’ll need
  • 6+ years engineering experience, 3+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience.
  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.).
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling.
  • Excellent problem-solving and troubleshooting skills.
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction.
  • Ability to partner deeply with product engineering teams
  • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)

Benefits

Comp & perks
  • Restricted Stock Units (RSUs)
  • 30 days annual leave
  • Grafana Shutdown Days to allow team to disconnect