Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Site Reliability Engineer – HPC

NVIDIA

Senior SRE building reliable HPC compute-farm services for NVIDIA’s accelerated-computing platform. Automating multi-cloud infrastructure, observability, and incident response.

Posted 9/4/2026full-timeSanta Clara • California, North Carolina, Texas • 🇺🇸 United StatesSenior💰 $152,000 - $287,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) solutions, including design, implementation, and operation within multi-cloud environments. Proficient in Infrastructure as Code, CI/CD techniques, and large-scale HPC cluster management, ensuring high availability and performance.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure as Code (IaC)High-Performance Computing (HPC) ClustersCI/CD TechniquesCoding/Scripting in Python, Go, Perl, or Ruby

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringInfrastructure as CodeCI/CD TechniquesHPC Cluster ManagementCapacity ManagementPerformance MonitoringCoding/Scripting in PythonCoding/Scripting in GoCoding/Scripting in PerlCoding/Scripting in Ruby
Soft Skills
Creative Problem-SolvingExcellent Debugging SkillsStrong CommunicationDocumentation Abilities
Tools & Technologies
SlurmLSFKubernetesAWSGCPOCIMonitoring ToolsMetrics ToolsContainer Management ToolsLog Collection Tools
Industry Keywords
Multi-Cloud EnvironmentE2E ObservabilityData-Driven OperationsAIOpsRoot-Cause Analysis (RCA)

Tech Stack

Tools & technologies
AWSCloudGoGoogle Cloud PlatformKubernetesPerlPythonRuby

About the role

Key responsibilities & impact
  • Own SRE solutions end-to-end, from design and implementation through operation and continuous improvement
  • Integrate SRE solutions with HPC schedulers, storage, and network fabrics
  • Use Infrastructure as Code and configuration management to standardize and automate provisioning
  • Deliver solutions in a globally distributed, multi-cloud hybrid environment across on-premises, AWS, GCP, and OCI
  • Design for failure using redundancy, failure domains, progressive delivery, and strict change control
  • Ensure uptime and Quality of Service for internal customers
  • Conduct capacity management and planning
  • Detect performance issues and recommend solutions
  • Collaborate with various teams to ensure seamless project completion
  • Participate in on-call rotations and incident reviews
  • Assist in root-cause identification and produce RCA reports

Requirements

What you’ll need
  • B.S. degree in Computer Science or related technical field (or equivalent experience)
  • 5+ years professional experience building and supporting critical services
  • Experience supporting large-scale HPC clusters using Slurm, LSF or Kubernetes, including setup, tuning, and troubleshooting
  • Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC)
  • Experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals)
  • Proficiency in monitoring, metrics, container management, and log collection tools
  • 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby
  • Experience mentoring engineers and influencing technical direction through design reviews, architecture documents, and partnership with product and leadership
  • Creative problem-solving, excellent debugging skills, and strong communication and documentation abilities

Benefits

Comp & perks
  • Equity
  • Comprehensive benefits package
  • Highly competitive salary