FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer
NVIDIASenior SRE improving NVIDIA GeForce NOW’s reliable GPU cloud gaming infrastructure. Building observability, automation, Kubernetes, and incident-response tooling for service SLOs.
Posted 8/11/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $168,000 - $270,250 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive Site Reliability Engineering expertise with a strong focus on Kubernetes, production automation, and incident management. Proficient in developing tools for observability and improving service reliability through effective communication and collaboration.
Highest-signal resume keywords
Site Reliability EngineeringKubernetes ManagementProduction AutomationIncident ResponseCloud Deployment Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringKubernetesGo ProgrammingPython ProgrammingBash ScriptingChange ManagementRoot-Cause AnalysisDeployment PipelinesAutomated Anomaly DetectionLog Clustering
Soft Skills
Problem-SolvingAnalytical SkillsCommunication SkillsPresentation SkillsSocial Skills
Tools & Technologies
DatadogPrometheusAlertmanagerGitHub ActionsGitLab CIArgoCDAWSGCPAzureAI Tools
Industry Keywords
MicroservicesService Level ObjectivesProduction SystemsOn-Call RotationVMI Setup
Tech Stack
Tools & technologiesAWSAzureCloudGoGoogle Cloud PlatformKubernetesMicroservicesPrometheusPython
About the role
Key responsibilities & impact- Build tools to improve SRE observability
- Participate in the Kubernetes migration journey with VMI setup and problem solving
- Rapidly debug and triage incidents and user-reported issues
- Automate, script, and develop tooling for new and existing scripts to achieve 100% automation of daily tasks
- Support services before launch through system design consulting, software platform and framework development, capacity management, and launch reviews
- Participate in an on-call rotation supporting production systems
- Drive tools and service development to maintain and improve service SLOs
- Partner with Service Owners to drive service reliability
- Lead production improvements, including change management, post-mortem reviews, workflow processes, and software automation
Requirements
What you’ll need- MS or BS in Computer Science, Engineering, or a related field, or equivalent experience
- 8+ years of Site Reliability Engineering experience with large-scale distributed microservices in production
- Strong Kubernetes background, including complex, highly available VMI setups
- Experience leading production improvements, change management, post-mortems, workflow processes, and software automation
- Problem-solving and root-cause analysis strengths
- Experience with Datadog, Prometheus, Alertmanager, or similar monitoring systems
- Experience managing multi-region cloud deployments on AWS, GCP, or Azure
- Experience designing and managing deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD
- Production-grade coding proficiency in Go, Python, or robust Bash scripting
- Required primary production on-call experience responding to and mitigating high-severity infrastructure alerts and service degradations
- Excellent communication, presentation, social, and analytical skills
- Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms is a plus
- Comfort using AI daily as an SRE
- Prior SRE or Service Engineer experience is a plus
Benefits
Comp & perks- Competitive salary package
- Equity
- Benefits