Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Grafana Labs

Senior Software Engineer – Databases, SRE

Grafana Labs

Senior Software Engineer - SRE supporting Grafana Cloud customer databases for exceptional reliability. Collaborating with engineering teams on production systems and automation for high-SLA customers.

Posted 7/20/2026full-timeRemote • 🇨🇦 CanadaSenior💰 CA$164,490 - CA$197,389 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on designing and implementing Service Level Objectives (SLOs) and automation for production reliability. Proficient in incident response, post-incident reviews, and collaborating with product engineering teams to enhance system scalability and operability.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Kubernetes ExperienceService Level Objectives (SLOs)Incident ResponseInfrastructure-as-Code Tooling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesAWSGCPAzureHelmTerraformJsonnetGoPythonJava
Soft Skills
Problem-SolvingTroubleshootingAutonomyCollaborationIntellectual Curiosity
Tools & Technologies
Linux Operating SystemsCloud StorageNetworkingMulti-Tenant Systems
Industry Keywords
Production EngineeringPost-Incident ReviewsSLO TargetsProduction ScalabilityOperational Reliability

Tech Stack

Tools & technologies
AWSAzureCloudGoGoogle Cloud PlatformJavaKubernetesLinuxPythonTerraform

About the role

Key responsibilities & impact
  • Partner closely with product engineering squads (embedded model)
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale our reliability practices
  • Ensuring our customers meet our SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serving as a primary escalation point and on-call for relevant incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews
  • Influence feature design to ensure production scalability and operability
  • Build automation to eliminate toil where needed
  • Improve alert quality and reduce noisy escalations

Requirements

What you’ll need
  • 6+ years engineering experience, 3+ in SRE/CRE/production engineering
  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)
  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
  • Ability to partner deeply with product engineering teams
  • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)

Benefits

Comp & perks
  • 100% Remote, Global Culture
  • Scaling Organization
  • Transparent Communication
  • Innovation-Driven
  • Open Source Roots
  • Empowered Teams
  • Career Growth Pathways
  • Approachable Leadership
  • Passionate People
  • In-Person onboarding
  • Balance is Key - 30 days annual leave, 3 days reserved for Grafana Shutdown Days