Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
GitLab

Site Reliability Engineer – Intermediate to Senior Staff, Infrastructure Platforms

GitLab

Site Reliability Engineer ensuring reliability of GitLab's user-facing services. Supporting operational excellence through engineering principles and automation.

Posted 7/11/2026full-timeRemote • 🌎 Anywhere in the WorldSenior💰 $126,400 - $314,400 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in maintaining reliable production systems and building automation tools using Infrastructure as Code. Proficient in Kubernetes operations, incident response, and observability practices to enhance system performance and reliability.

Highest-signal resume keywords
Kubernetes OperationsInfrastructure As CodeTerraform ModulesCloud Provider Experience (GCP or AWS)Observability Practices

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Go ProgrammingRuby ProgrammingCI/CDGitOpsProduction AutomationDebuggingMetricsLoggingSLOsIncident Response
Soft Skills
Strong Written CommunicationTroubleshooting Under PressureAsync Collaboration
Industry Keywords
Production SystemsAutomationInfrastructure ToolingObservability StackRunbooks

Tech Stack

Tools & technologies
AWSCloudGoGoogle Cloud PlatformKubernetesRubyTerraform

About the role

Key responsibilities & impact
  • Keep user-facing services and production systems reliable, scalable, and efficient
  • Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
  • Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps
  • Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately
  • Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
  • Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
  • Document runbooks, architecture decisions, and reviews so your findings become repeatable practices

Requirements

What you’ll need
  • Experience keeping production systems reliable, combining an operations mindset with real software engineering practice
  • Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
  • The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes
  • Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level
  • Hands-on experience with at least one major cloud provider (GCP or AWS)
  • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions
  • Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure
  • Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment
  • A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work
  • Alignment with GitLab's values and a commitment to working in accordance with them.

Benefits

Comp & perks
  • Benefits to support your health, finances, and well-being
  • Flexible Paid Time Off
  • Team Member Resource Groups
  • Equity Compensation & Employee Stock Purchase Plan
  • Growth and Development Fund
  • Parental Leave