Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Engineering Manager, Deep Learning Inference

NVIDIA

Engineering Manager leading a team focusing on deep learning inference and GPU-accelerated software. Guiding strategy and execution of open-source inference frameworks for advanced AI systems.

Posted 7/29/2026full-timeRemote • California, District of Columbia, Illinois • 🇺🇸 United StatesMid-LevelSenior💰 $224,000 - $431,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in leading and mentoring engineering teams focused on deep learning inference and GPU-accelerated software, with a strong emphasis on performance optimization and best practices in CUDA and multi-GPU communications.

Highest-signal resume keywords
Technical LeadershipDeep Learning OptimizationGPU Programming (CUDA, Triton, CUTLASS)Agile Software DevelopmentPerformance Tuning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
C/C++ Software DesignPython ProgrammingDeep Learning Model DeploymentPerformance OptimizationSoftware Development Experience
Soft Skills
MentoringCollaborationStrategic PlanningInnovation
Tools & Technologies
CUDATritonCUTLASSNIXLNCCLNVSHMEM
Certifications & Qualifications
MS in Computer SciencePhD in Computer ScienceEquivalent Experience
Industry Keywords
Deep Learning InferenceGPU-Accelerated SoftwareLarge-Scale ModelsMultimodal AIGenerative AI

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software
  • Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering
  • Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM)
  • Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies
  • Foster a culture of technical excellence, open collaboration, and continuous innovation

Requirements

What you’ll need
  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field
  • 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management
  • Strong background in C/C++ software design and development; proficiency in Python is a plus
  • Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization
  • Proven record of deploying or optimizing deep learning models in production environments
  • Experience leading teams using Agile or collaborative software development practices

Benefits

Comp & perks
  • highly competitive salaries
  • comprehensive benefits package