Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior System Software Engineer, Agentic Inference - Dynamo

NVIDIA

Senior System Software Engineer developing open-source software for AI model inference at NVIDIA. Contributing to distributed serving and optimizing performance with a focus on agentic workloads and modern LLMs.

Posted 7/28/2026full-timeSanta Clara • California • 🇺🇸 United StatesSenior💰 $224,000 - $431,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing open source software for AI model inference on GPUs, with a strong focus on building scalable and high-performance distributed systems. Proficient in Rust and Python programming, with a deep understanding of modern LLM API semantics and inference-state management.

Highest-signal resume keywords
Rust ProgrammingPython ProgrammingDistributed Inference SystemsLLM API SemanticsAgile Team Collaboration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software DesignDebuggingPerformance AnalysisTest DesignInference-State Management
Soft Skills
CollaborationAdaptability
Tools & Technologies
DynamoVLLMSGLangTensorRT-LLM
Certifications & Qualifications
Masters DegreePhD
Industry Keywords
AI ModelsInference WorkloadsOpen Source SoftwareComputer ScienceComputer Engineering

Tech Stack

Tools & technologies
Open SourcePythonRust

About the role

Key responsibilities & impact
  • develop open source software to serve inference of trained AI models running on GPUs
  • contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads
  • innovate in inference-state management for long-running agents
  • build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM
  • balance a variety of objectives: build robust, scalable, high performance software components to support our distributed inference workloads

Requirements

What you’ll need
  • Masters or PhD or equivalent experience
  • 10+ years in Computer Science, Computer Engineering, or related field
  • Ability to work in a fast-paced, agile team environment
  • Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design
  • Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs

Benefits

Comp & perks
  • equity
  • benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score