Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA

Senior software engineer researching and optimizing CUDA kernels, distributed systems, and deep-learning models at NVIDIA. Building high-performance AI infrastructure from single GPUs to supercomputer clusters.

Posted 8/5/2026full-timeRemote • California, Texas • 🇺🇸 United StatesSenior💰 $184,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in deep learning model optimization, CUDA programming, and distributed computing systems. Proficient in analyzing hardware-software interactions and developing high-performance solutions for advanced AI applications.

Highest-signal resume keywords
C++ ProgrammingPython ProgrammingCUDA ProgrammingDeep Learning FundamentalsDistributed Computing

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDA Kernel OptimizationPerformance Bottleneck AnalysisSystems ProgrammingComputer ArchitectureWorkload ProfilingGenerative AI Model OptimizationTransformersMulti-Node ScalingLow-Level Systems Performance OptimizationDeep Learning Compilers
Soft Skills
CollaborationInitiativeProblem-Solving
Tools & Technologies
PyTorchJAXTensorRTNCCLMPIUCXTritonXLATorch.compileSgLang
Industry Keywords
Deep LearningAI SystemsNeural Network ArchitecturesCluster-Scale PerformanceGenerative AI ArchitecturesVision ModelsDiffusion ModelsNumerical MethodsLow-Precision ArithmeticAgentic AI Systems

Tech Stack

Tools & technologies
Node.jsPythonPyTorch

About the role

Key responsibilities & impact
  • Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping
  • Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms
  • Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products

Requirements

What you’ll need
  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience
  • 8+ years of relevant industry experience or equivalent academic experience after degree achievement
  • Strong proficiency in C++ and Python programming
  • Solid background in deep learning fundamentals, with a focus on transformers
  • Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization
  • Familiarity with GPU accelerator architectures
  • Hands-on experience with CUDA programming, kernel optimization, and workload profiling
  • Experience profiling and optimizing generative AI models, including large language models
  • Research background in machine learning systems or adjacent fields
  • Experience profiling and optimizing vision models, generative AI architectures, or diffusion models
  • Track record of initiative and willingness to deep-dive on problems across the stack
  • Preferred: expertise in performance internals and execution graphs of PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron
  • Preferred: experience with NCCL, MPI, UCX, and distributed machine learning techniques such as pipeline, tensor, or expert parallelism
  • Preferred: knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, or INT8
  • Preferred: background in deep learning compilers and ML systems, including Triton, XLA, or torch.compile
  • Preferred: experience designing agentic AI systems for complex systems and infrastructure problems

Benefits

Comp & perks
  • Equity
  • Benefits
  • Equal opportunity employer
  • Inclusive work environment