FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer – HPC
NVIDIASenior SRE building reliable HPC compute-farm services for NVIDIA’s accelerated-computing platform. Automating multi-cloud infrastructure, observability, and incident response.
Posted 9/4/2026full-timeSanta Clara • California, North Carolina, Texas • 🇺🇸 United StatesSenior💰 $152,000 - $287,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) solutions, including design, implementation, and operation within multi-cloud environments. Proficient in Infrastructure as Code, CI/CD techniques, and large-scale HPC cluster management, ensuring high availability and performance.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure as Code (IaC)High-Performance Computing (HPC) ClustersCI/CD TechniquesCoding/Scripting in Python, Go, Perl, or Ruby
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringInfrastructure as CodeCI/CD TechniquesHPC Cluster ManagementCapacity ManagementPerformance MonitoringCoding/Scripting in PythonCoding/Scripting in GoCoding/Scripting in PerlCoding/Scripting in Ruby
Soft Skills
Creative Problem-SolvingExcellent Debugging SkillsStrong CommunicationDocumentation Abilities
Tools & Technologies
SlurmLSFKubernetesAWSGCPOCIMonitoring ToolsMetrics ToolsContainer Management ToolsLog Collection Tools
Industry Keywords
Multi-Cloud EnvironmentE2E ObservabilityData-Driven OperationsAIOpsRoot-Cause Analysis (RCA)
Tech Stack
Tools & technologiesAWSCloudGoGoogle Cloud PlatformKubernetesPerlPythonRuby
About the role
Key responsibilities & impact- Own SRE solutions end-to-end, from design and implementation through operation and continuous improvement
- Integrate SRE solutions with HPC schedulers, storage, and network fabrics
- Use Infrastructure as Code and configuration management to standardize and automate provisioning
- Deliver solutions in a globally distributed, multi-cloud hybrid environment across on-premises, AWS, GCP, and OCI
- Design for failure using redundancy, failure domains, progressive delivery, and strict change control
- Ensure uptime and Quality of Service for internal customers
- Conduct capacity management and planning
- Detect performance issues and recommend solutions
- Collaborate with various teams to ensure seamless project completion
- Participate in on-call rotations and incident reviews
- Assist in root-cause identification and produce RCA reports
Requirements
What you’ll need- B.S. degree in Computer Science or related technical field (or equivalent experience)
- 5+ years professional experience building and supporting critical services
- Experience supporting large-scale HPC clusters using Slurm, LSF or Kubernetes, including setup, tuning, and troubleshooting
- Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC)
- Experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals)
- Proficiency in monitoring, metrics, container management, and log collection tools
- 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby
- Experience mentoring engineers and influencing technical direction through design reviews, architecture documents, and partnership with product and leadership
- Creative problem-solving, excellent debugging skills, and strong communication and documentation abilities
Benefits
Comp & perks- Equity
- Comprehensive benefits package
- Highly competitive salary