FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer – Databases, SRE
Grafana LabsSenior Software Engineer focusing on customer reliability for Grafana Cloud databases. Designing solutions for SLOs and incident response, ensuring high reliability for complex environments.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on designing and implementing Service Level Objectives (SLOs) and automation for production reliability. Proficient in Kubernetes, cloud environments, and incident response processes, ensuring high availability and performance in complex customer systems.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Kubernetes ExperienceService Level Objectives (SLOs) DesignIncident Response ParticipationAutomation Implementation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesService Level Objectives (SLOs)GoPythonJavaLinux InternalsInfrastructure-as-CodeTerraformHelmJsonnet
Soft Skills
Problem-SolvingTroubleshootingIntellectual CuriosityTransparencyKindness
Tools & Technologies
AWSGCPAzureCloud StorageNetworking
Industry Keywords
Production EngineeringMulti-Tenant SystemsPost Incident Reviews (PIRs)Customer Reliability Engineering
Tech Stack
Tools & technologiesAWSAzureCloudGoGoogle Cloud PlatformJavaKubernetesLinuxPythonTerraform
About the role
Key responsibilities & impact- Partner closely with product engineering squads (embedded model)
- Own production reliability for high-SLA and complex customer environments
- Design and implement automation to scale our reliability practices
- Ensuring our customers meet our SLO targets
- Define and evolve per-tenant SLOs and reliability models
- Proactively reduce SLO burn to prevent repeat incidents
- Serving as a primary escalation point and on-call for relevant incidents
- Lead customer-impacting incident response and post-incident reviews
- Contribute to design docs and code reviews
- Influence feature design to ensure production scalability and operability
- Build automation to eliminate toil where needed
- Improve alert quality and reduce noisy escalations
Requirements
What you’ll need- 6+ years engineering experience, 3+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience.
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.).
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling.
- Excellent problem-solving and troubleshooting skills.
- Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction.
- Ability to partner deeply with product engineering teams
- We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
Benefits
Comp & perks- Restricted Stock Units (RSUs)
- 30 days annual leave
- Grafana Shutdown Days to allow team to disconnect