FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Software Engineer, CUDA Deep Learning Systems
NVIDIANVIDIA software engineer researching CUDA kernels and distributed systems for advanced deep-learning workloads. Optimizing GPU performance across training, inference, and cluster-scale AI infrastructure.
Posted 8/5/2026full-timeSanta Clara • California, Texas • 🇺🇸 United StatesJuniorMid-Level💰 $124,000 - $195,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in CUDA programming, high-performance computing, and deep learning optimization, with a strong foundation in distributed systems and machine learning architectures. Capable of collaborating with cross-functional teams to design and implement innovative solutions for advanced AI models.
Highest-signal resume keywords
CUDA ProgrammingC++ and Python ProficiencyDeep Learning OptimizationDistributed Computing SystemsPerformance Bottleneck Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
CUDAC++PythonDeep Learning FundamentalsSystems ProgrammingComputer ArchitectureKernel OptimizationWorkload ProfilingGenerative AI ModelsTransformers
Soft Skills
CollaborationInitiativeProblem-Solving
Tools & Technologies
PyTorchJAXTensorRTNCCLMPIUCXTritonXLATorch.compileSgLang
Industry Keywords
Machine Learning SystemsHigh-Performance ComputingAI ResearchNeural Network ArchitecturesCluster-Scale Performance
Tech Stack
Tools & technologiesNode.jsPythonPyTorch
About the role
Key responsibilities & impact- Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA.
- Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments.
- Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
- Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
- Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
- Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
- Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products.
Requirements
What you’ll need- BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
- 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
- Strong proficiency in C++ and Python programming.
- Solid background in deep learning fundamentals, focused on transformers.
- Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges.
- Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
- Familiarity with GPU deep learning accelerator architectures.
- Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
- Experience profiling and optimizing generative AI models, including large language models.
- Research background in machine learning systems or adjacent fields.
- Experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
- Track record of initiative and willingness to deep-dive on problems across the stack.
- Preferred experience with PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron internals and execution graphs.
- Preferred hands-on experience with NCCL, MPI, or UCX and distributed machine learning techniques.
- Preferred knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, and INT8.
- Preferred background in deep learning compilers and ML systems, including Triton, XLA, and torch.compile.
- Preferred experience with highly parallel or reinforcement-learning-style simulation environments.
- Preferred experience designing and implementing agentic AI systems for complex systems and infrastructure problems.
Benefits
Comp & perks- Equity
- Benefits
- Equal opportunity employer
- Inclusive work environment