Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Deep Learning Software Engineer, Inference

NVIDIA

Deep Learning Software Engineer specializing in inference for NVIDIA. Focusing on performance optimizations across AI applications and frameworks.

Posted 7/23/2026full-timeRemote • California, Massachusetts, New York, Texas, Washington • 🇺🇸 United StatesMid-LevelSenior💰 $124,000 - $195,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in performance optimization and tuning of Deep Learning models, with strong programming skills in C/C++ and experience in GPU programming. Familiarity with NVIDIA libraries and frameworks, along with a solid foundation in software development and Agile methodologies, is essential.

Highest-signal resume keywords
C/C++ ProgrammingDeep Learning Model OptimizationGPU Programming (CUDA, OAI TRITON, CUTLASS)Software Development ExperiencePerformance Modeling and Profiling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Performance OptimizationDeep LearningSoftware DesignCode OptimizationArchitectural Knowledge of CPU and GPUInference OptimizationAgile MethodologiesPython ProgrammingModel Training and DeploymentDebugging
Tools & Technologies
NVIDIA AcceleratorsVLLMSGLangFlashInferLLM Software Solutions
Industry Keywords
Machine LearningArtificial IntelligenceMultimodal AIGenerative AICross-Collaborative Teams

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI.
  • Scale performance of DL models across different architectures and types of NVIDIA accelerators.
  • Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions.
  • Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions.

Requirements

What you’ll need
  • Pursuing or recently completed a MS or PhD Computer Engineering, Computer Science, EECS, AI or related field or equivalent experience.
  • Software development experience.
  • Excellent C/C++ programming and software design skills.
  • SW Agile skills are helpful and Python experience is a plus.
  • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus.
  • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus.
  • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus.

Benefits

Comp & perks
  • equity
  • benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score