FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Engineering Manager, Deep Learning Inference
NVIDIAEngineering Manager leading a team focusing on deep learning inference and GPU-accelerated software. Guiding strategy and execution of open-source inference frameworks for advanced AI systems.
Posted 7/29/2026full-timeRemote • California, District of Columbia, Illinois • 🇺🇸 United StatesMid-LevelSenior💰 $224,000 - $431,250 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in leading and mentoring engineering teams focused on deep learning inference and GPU-accelerated software, with a strong emphasis on performance optimization and best practices in CUDA and multi-GPU communications.
Highest-signal resume keywords
Technical LeadershipDeep Learning OptimizationGPU Programming (CUDA, Triton, CUTLASS)Agile Software DevelopmentPerformance Tuning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
C/C++ Software DesignPython ProgrammingDeep Learning Model DeploymentPerformance OptimizationSoftware Development Experience
Soft Skills
MentoringCollaborationStrategic PlanningInnovation
Tools & Technologies
CUDATritonCUTLASSNIXLNCCLNVSHMEM
Certifications & Qualifications
MS in Computer SciencePhD in Computer ScienceEquivalent Experience
Industry Keywords
Deep Learning InferenceGPU-Accelerated SoftwareLarge-Scale ModelsMultimodal AIGenerative AI
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software
- Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering
- Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators
- Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications
- Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM)
- Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies
- Foster a culture of technical excellence, open collaboration, and continuous innovation
Requirements
What you’ll need- MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field
- 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management
- Strong background in C/C++ software design and development; proficiency in Python is a plus
- Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization
- Proven record of deploying or optimizing deep learning models in production environments
- Experience leading teams using Agile or collaborative software development practices
Benefits
Comp & perks- highly competitive salaries
- comprehensive benefits package