FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer – Databases, SRE
Grafana LabsSenior Software Engineer - SRE supporting Grafana Cloud customer databases for exceptional reliability. Collaborating with engineering teams on production systems and automation for high-SLA customers.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on designing and implementing Service Level Objectives (SLOs) and automation for production reliability. Proficient in incident response, post-incident reviews, and collaborating with product engineering teams to enhance system scalability and operability.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Kubernetes ExperienceService Level Objectives (SLOs)Incident ResponseInfrastructure-as-Code Tooling
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesAWSGCPAzureHelmTerraformJsonnetGoPythonJava
Soft Skills
Problem-SolvingTroubleshootingAutonomyCollaborationIntellectual Curiosity
Tools & Technologies
Linux Operating SystemsCloud StorageNetworkingMulti-Tenant Systems
Industry Keywords
Production EngineeringPost-Incident ReviewsSLO TargetsProduction ScalabilityOperational Reliability
Tech Stack
Tools & technologiesAWSAzureCloudGoGoogle Cloud PlatformJavaKubernetesLinuxPythonTerraform
About the role
Key responsibilities & impact- Partner closely with product engineering squads (embedded model)
- Own production reliability for high-SLA and complex customer environments
- Design and implement automation to scale our reliability practices
- Ensuring our customers meet our SLO targets
- Define and evolve per-tenant SLOs and reliability models
- Proactively reduce SLO burn to prevent repeat incidents
- Serving as a primary escalation point and on-call for relevant incidents
- Lead customer-impacting incident response and post-incident reviews
- Contribute to design docs and code reviews
- Influence feature design to ensure production scalability and operability
- Build automation to eliminate toil where needed
- Improve alert quality and reduce noisy escalations
Requirements
What you’ll need- 6+ years engineering experience, 3+ in SRE/CRE/production engineering
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting skills
- Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
- Ability to partner deeply with product engineering teams
- We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
Benefits
Comp & perks- 100% Remote, Global Culture
- Scaling Organization
- Transparent Communication
- Innovation-Driven
- Open Source Roots
- Empowered Teams
- Career Growth Pathways
- Approachable Leadership
- Passionate People
- In-Person onboarding
- Balance is Key - 30 days annual leave, 3 days reserved for Grafana Shutdown Days