FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer, Resilience Engineering - DGX Cloud
NVIDIASenior Software Engineer focusing on resilience engineering for NVIDIA's DGX Cloud. Building reliability strategies and leading incident responses to enhance operational excellence.
Posted 7/28/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $184,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and implementing reliability strategies, including SLO programs and chaos engineering practices, while leading incident response and enhancing operational standards. Possesses strong software engineering skills in Go and Python, with a focus on improving production code and tooling.
Highest-signal resume keywords
SLO Program DevelopmentChaos EngineeringFailure InjectionGo ProgrammingPython Programming
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringOperational PracticesProduction Code ImprovementIncident ResponseResilience Testing
Soft Skills
Influencing Across TeamsLeadership
Certifications & Qualifications
Bachelor's DegreeMaster's Degree
Industry Keywords
Reliability Strategy24/7 EnvironmentHigh Severity IncidentsOperational Rigor
Tech Stack
Tools & technologiesGoPython
About the role
Key responsibilities & impact- Build org-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment
- Stand up a rigorous SLO program, defining and maintaining high standards across teams
- Lead incident response for high severity incidents, ensuring low drama and high signal resolution
- Build and improve production code daily, enhancing our data platform and related tooling
- Implement chaos engineering, failure injection, and resilience testing to elevate our team's standard practices
- Improve standards by setting an example with your hands-on experience and leadership
Requirements
What you’ll need- 8+ years of industry experience
- Bachelor's or Master's degree, or equivalent experience operating systems at scale
- Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages
- Proven experience in establishing and maintaining an SLO program with operational rigor
- Practical experience in reliability fields such as chaos engineering and failure injection
- Ability to influence across team boundaries through credibility and expertise.
Benefits
Comp & perks- Equity
- Comprehensive benefits package