Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Software Engineer, Resilience Engineering - DGX Cloud

NVIDIA

Senior Software Engineer focusing on resilience engineering for NVIDIA's DGX Cloud. Building reliability strategies and leading incident responses to enhance operational excellence.

Posted 7/28/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $184,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and implementing reliability strategies, including SLO programs and chaos engineering practices, while leading incident response and enhancing operational standards. Possesses strong software engineering skills in Go and Python, with a focus on improving production code and tooling.

Highest-signal resume keywords
SLO Program DevelopmentChaos EngineeringFailure InjectionGo ProgrammingPython Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringOperational PracticesProduction Code ImprovementIncident ResponseResilience Testing
Soft Skills
Influencing Across TeamsLeadership
Certifications & Qualifications
Bachelor's DegreeMaster's Degree
Industry Keywords
Reliability Strategy24/7 EnvironmentHigh Severity IncidentsOperational Rigor

Tech Stack

Tools & technologies
GoPython

About the role

Key responsibilities & impact
  • Build org-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment
  • Stand up a rigorous SLO program, defining and maintaining high standards across teams
  • Lead incident response for high severity incidents, ensuring low drama and high signal resolution
  • Build and improve production code daily, enhancing our data platform and related tooling
  • Implement chaos engineering, failure injection, and resilience testing to elevate our team's standard practices
  • Improve standards by setting an example with your hands-on experience and leadership

Requirements

What you’ll need
  • 8+ years of industry experience
  • Bachelor's or Master's degree, or equivalent experience operating systems at scale
  • Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages
  • Proven experience in establishing and maintaining an SLO program with operational rigor
  • Practical experience in reliability fields such as chaos engineering and failure injection
  • Ability to influence across team boundaries through credibility and expertise.

Benefits

Comp & perks
  • Equity
  • Comprehensive benefits package