Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Calix

Staff Site Reliability Operations Engineer

Calix

Staff Site Reliability Engineer leading global platform reliability and observability strategy at Calix, specializing in GCP and Kubernetes infrastructure.

Posted 7/29/2026full-timeRemote • 🇺🇸 United StatesLead💰 $136,000 - $231,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in leading platform reliability and observability strategies on GCP, with a strong focus on architecting and optimizing complex networking infrastructures and observability platforms. Proficient in deploying machine learning models and automating incident response while mentoring engineers in advanced debugging techniques.

Highest-signal resume keywords
Google Kubernetes Engine (GKE)Grafana Enterprise/CloudHashiCorp TerraformGo ProgrammingPython Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability Engineering (SRE)Production EngineeringDistributed SystemsNetworking InfrastructureAnomaly DetectionDatabase Performance OptimizationIncident Response AutomationDebugging TechniquesApache Kafka ManagementOSI Layer Knowledge
Soft Skills
Mentoring
Tools & Technologies
GrafanaPrometheusMimirLokiTempoPostgreSQLAlloyDBBigQueryAIOps

Tech Stack

Tools & technologies
ApacheBigQueryCloudDistributed SystemsGoGoogle Cloud PlatformGrafanaKafkaKubernetesPostgresPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Lead global platform reliability and observability strategy on GCP
  • Architect and troubleshoot complex networking infrastructure
  • Design and optimize observability platform using Grafana
  • Deploy machine learning models for anomaly detection
  • Drive architecture and security of GKE clusters
  • Ensure performance and scalability of databases (PostgreSQL, AlloyDB, BigQuery)
  • Automate incident response and triaging using AIOps
  • Mentor engineers on advanced debugging techniques

Requirements

What you’ll need
  • 8+ years in SRE, Production Engineering, or Distributed Systems
  • Deep technical knowledge across OSI layers (L1-L7)
  • Expert-level mastery of Google Kubernetes Engine (GKE)
  • Proven track record managing Apache Kafka pipelines and large-scale data environments
  • Experience deploying and managing Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
  • Advanced expertise utilizing HashiCorp Terraform for GCP architectures
  • High proficiency in Go and Python

Benefits

Comp & perks
  • Health insurance
  • Bonuses