FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in leading platform reliability and observability strategies on GCP, with a strong focus on architecting and optimizing complex networking infrastructures and observability platforms. Proficient in deploying machine learning models and automating incident response while mentoring engineers in advanced debugging techniques.
Highest-signal resume keywords
Google Kubernetes Engine (GKE)Grafana Enterprise/CloudHashiCorp TerraformGo ProgrammingPython Programming
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability Engineering (SRE)Production EngineeringDistributed SystemsNetworking InfrastructureAnomaly DetectionDatabase Performance OptimizationIncident Response AutomationDebugging TechniquesApache Kafka ManagementOSI Layer Knowledge
Soft Skills
Mentoring
Tools & Technologies
GrafanaPrometheusMimirLokiTempoPostgreSQLAlloyDBBigQueryAIOps
Tech Stack
Tools & technologiesApacheBigQueryCloudDistributed SystemsGoGoogle Cloud PlatformGrafanaKafkaKubernetesPostgresPrometheusPythonTerraform
About the role
Key responsibilities & impact- Lead global platform reliability and observability strategy on GCP
- Architect and troubleshoot complex networking infrastructure
- Design and optimize observability platform using Grafana
- Deploy machine learning models for anomaly detection
- Drive architecture and security of GKE clusters
- Ensure performance and scalability of databases (PostgreSQL, AlloyDB, BigQuery)
- Automate incident response and triaging using AIOps
- Mentor engineers on advanced debugging techniques
Requirements
What you’ll need- 8+ years in SRE, Production Engineering, or Distributed Systems
- Deep technical knowledge across OSI layers (L1-L7)
- Expert-level mastery of Google Kubernetes Engine (GKE)
- Proven track record managing Apache Kafka pipelines and large-scale data environments
- Experience deploying and managing Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
- Advanced expertise utilizing HashiCorp Terraform for GCP architectures
- High proficiency in Go and Python
Benefits
Comp & perks- Health insurance
- Bonuses
