Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Zscaler

Production Engineer

Zscaler

Production Engineer at Zscaler ensuring reliability and scalability of a cloud-native platform processing billions of transactions daily. Collaborates across teams while driving an automation-first culture.

Posted 6/30/2026full-timeSan Jose • California • 🇺🇸 United StatesJunior💰 $102,400 - $128,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in implementing scalable infrastructure across AWS and GCP, with a strong focus on automation through programming in Python and Go. Proven ability to manage incident response and enhance service reliability using observability tools like Prometheus and Grafana.

Highest-signal resume keywords
AWS Infrastructure ManagementGCP Infrastructure ManagementPython ProgrammingIncident ManagementObservability Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonGoC/C++Networking ProtocolsLinux/RHEL SystemsDistributed ArchitectureIncident ResponseService ReliabilityAutomationProblem Management
Soft Skills
CuriosityCollaborationAnalytical ThinkingCommunication
Tools & Technologies
PrometheusGrafanaOpenTelemetryITIL Frameworks
Industry Keywords
ScalabilityAvailabilitySelf-Healing SystemsError BudgetsSLIs/SLOs

Tech Stack

Tools & technologies
AWSGoGoogle Cloud PlatformGrafanaLinuxPrometheusPython

About the role

Key responsibilities & impact
  • Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments
  • Drive an "automation-first" culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
  • Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets
  • Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses
  • Partner with Engineering and partner teams to conduct operability reviews

Requirements

What you’ll need
  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • 1-3 years of experience managing reliability, scalability, and availability for large-scale production services
  • Deep expertise in programming (e.g., Python, Go, or C/C++)
  • Strong background in networking protocols, Linux/RHEL systems, and distributed architecture
  • Experience in high-stakes incident management and participation in a 24/7 on-call rotation
  • Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews

Benefits

Comp & perks
  • Various health plans
  • Time off plans for vacation and sick time
  • Parental leave options
  • Retirement options
  • Education reimbursement
  • In-office perks, and more!