Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Software Engineer, CUDA Deep Learning Systems

NVIDIA

NVIDIA software engineer researching CUDA kernels and distributed systems for advanced deep-learning workloads. Optimizing GPU performance across training, inference, and cluster-scale AI infrastructure.

Posted 8/5/2026full-timeSanta Clara • California, Texas • 🇺🇸 United StatesJuniorMid-Level💰 $124,000 - $195,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in CUDA programming, high-performance computing, and deep learning optimization, with a strong foundation in distributed systems and machine learning architectures. Capable of collaborating with cross-functional teams to design and implement innovative solutions for advanced AI models.

Highest-signal resume keywords
CUDA ProgrammingC++ and Python ProficiencyDeep Learning OptimizationDistributed Computing SystemsPerformance Bottleneck Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDAC++PythonDeep Learning FundamentalsSystems ProgrammingComputer ArchitectureKernel OptimizationWorkload ProfilingGenerative AI ModelsTransformers
Soft Skills
CollaborationInitiativeProblem-Solving
Tools & Technologies
PyTorchJAXTensorRTNCCLMPIUCXTritonXLATorch.compileSgLang
Industry Keywords
Machine Learning SystemsHigh-Performance ComputingAI ResearchNeural Network ArchitecturesCluster-Scale Performance

Tech Stack

Tools & technologies
Node.jsPythonPyTorch

About the role

Key responsibilities & impact
  • Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA.
  • Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products.

Requirements

What you’ll need
  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
  • 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python programming.
  • Solid background in deep learning fundamentals, focused on transformers.
  • Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with GPU deep learning accelerator architectures.
  • Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
  • Experience profiling and optimizing generative AI models, including large language models.
  • Research background in machine learning systems or adjacent fields.
  • Experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
  • Track record of initiative and willingness to deep-dive on problems across the stack.
  • Preferred experience with PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron internals and execution graphs.
  • Preferred hands-on experience with NCCL, MPI, or UCX and distributed machine learning techniques.
  • Preferred knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, and INT8.
  • Preferred background in deep learning compilers and ML systems, including Triton, XLA, and torch.compile.
  • Preferred experience with highly parallel or reinforcement-learning-style simulation environments.
  • Preferred experience designing and implementing agentic AI systems for complex systems and infrastructure problems.

Benefits

Comp & perks
  • Equity
  • Benefits
  • Equal opportunity employer
  • Inclusive work environment