FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Engineering Manager, Deep Learning Inference
NVIDIAEngineering Manager leading NVIDIA’s GPU-accelerated deep learning inference software team. Driving frameworks, model optimization, and distributed inference for advanced AI systems.
Posted 8/5/2026full-timeSanta Clara • California, District of Columbia, New York, Texas, Washington • 🇺🇸 United StatesMid-LevelSenior💰 $184,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in leading and mentoring engineering teams focused on deep learning inference and GPU-accelerated software, with a strong emphasis on performance optimization and deployment of large-scale models. Proficient in CUDA, Triton, and CUTLASS, with a solid foundation in C/C++ and Python programming.
Highest-signal resume keywords
Technical LeadershipDeep Learning Model OptimizationCUDA ProgrammingPerformance TuningAgile Software Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
C/C++ Software DesignPython ProgrammingGPU ProgrammingPerformance OptimizationDeep Learning FrameworksInference FrameworksModel DeploymentSystem-Level OptimizationPerformance ProfilingArchitectural Decision-Making
Soft Skills
MentoringCollaborationInnovationCommunicationTeam Leadership
Tools & Technologies
CUDATritonCUTLASSNIXLNCCLNVSHMEMAgile MethodologiesDistributed Inference Architectures
Industry Keywords
Deep LearningGPU AccelerationInference PipelinesLarge-Scale ModelsMultimodal AIGenerative AIOpen-Source Contributions
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software
- Drive the strategy, roadmap, and execution of NVIDIA’s inference frameworks engineering, focusing on Client AI
- Partner with internal compiler, libraries, and research teams to deliver optimized inference pipelines across NVIDIA accelerators
- Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications
- Guide engineers in adopting CUDA, Triton, CUTLASS, and multi-GPU communication technologies
- Represent the team in roadmap and planning discussions
- Foster technical excellence, collaboration, and continuous innovation
Requirements
What you’ll need- MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field
- 6+ overall years of software development experience
- 3+ years in technical leadership or engineering management
- Strong background in C/C++ software design and development
- Proficiency in Python is a plus
- Hands-on experience with GPU programming using CUDA, Triton, and CUTLASS
- Experience with performance optimization
- Proven record of deploying or optimizing deep learning models in production environments
- Experience leading teams using Agile or collaborative software development practices
- Significant open-source contributions to deep learning or inference frameworks are advantageous
- Deep understanding of NIXL, NCCL, NVSHMEM, and distributed inference architectures is advantageous
- Expertise in performance modeling, profiling, and system-level optimization is advantageous
- Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects is advantageous
- Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering are advantageous
Benefits
Comp & perks- Competitive salary
- Equity
- Comprehensive benefits package
- Career advancement opportunities
- Inclusive work environment
- Equal opportunity employment