FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer – Distributed Systems
NVIDIASenior Software Engineer developing AI infrastructure solutions at NVIDIA. Designing systems for large scalable GPU clusters to support various AI workloads.
Posted 7/28/2026full-timeRemote • California, Texas, Washington • 🇺🇸 United StatesSenior💰 $152,000 - $287,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and developing scalable platforms for AI workloads, with a strong foundation in systems programming and incident management. Proven ability to collaborate effectively across teams and improve production systems for optimal performance.
Highest-signal resume keywords
Systems Programming (Go, Python)Large-Scale Production SystemsIncident Management ProcessData Structures and AlgorithmsStrong Communication Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Systems ProgrammingData StructuresAlgorithmsSoftware EngineeringIncident ManagementAI WorkloadsScalable PlatformsPerformance OptimizationProduction SystemsTechnical Impact
Soft Skills
Highly MotivatedEffective CoordinationTeam CollaborationCommunication Skills
Certifications & Qualifications
BS in Computer ScienceBS in EngineeringBS in PhysicsBS in Mathematics
Industry Keywords
GPU AssetsAI ClustersNVIDIAMulti-Functional Teams
Tech Stack
Tools & technologiesCloudGoPython
About the role
Key responsibilities & impact- Designing and developing a massively distributed scalable platform used to identify, diagnose and remediate non-performant GPU assets
- Working with teams across NVIDIA to ensure production AI clusters run reliably and consistently with maximum performance
- Evaluating system failures and improving services based on a well-defined incident management process
- Be part of a DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads
Requirements
What you’ll need- 5+ years in similar role and experience on large-scale production systems
- Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work
- Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms
- BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience
- Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies.
Benefits
Comp & perks- equity
- benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score