Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

AI Infrastructure Software Engineer

NVIDIA

AI infrastructure engineer building scalable pre-training, inference, and reinforcement-learning systems for NVIDIA’s Physical AI foundation models. Optimizing distributed GPU training, simulation integration, reliability, and performance.

Posted 8/12/2026full-timeShanghai • 🇨🇳 ChinaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and implementing training infrastructure for AI models, with a strong focus on reinforcement learning and distributed systems. Proficient in Python, C/C++/CUDA, and deep learning frameworks, with a proven ability to optimize performance and scalability in large-scale environments.

Highest-signal resume keywords
Python ProficiencyDistributed Systems DevelopmentReinforcement Learning InfrastructureDeep Learning Frameworks (PyTorch)Debugging and Triage Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software Infrastructure DevelopmentLarge-Scale Data LoadingCheckpointingInference Engine DevelopmentMixed Precision OptimizationAsynchronous Reinforcement LearningProduction-Grade InfrastructureTesting and Defensive ProgrammingVersion ControlCI Practices
Tools & Technologies
PyTorch (FSDP/DTensor)MegatronCUDASimulation EnvironmentsVectorized Environments
Certifications & Qualifications
Bachelor's Degree in Computer Science or Related Field
Industry Keywords
AI TrainingDistributed TrainingInference ServingSimulation-Robotics IntegrationReliability Metrics

Tech Stack

Tools & technologies
Distributed SystemsPythonPyTorch

About the role

Key responsibilities & impact
  • Create and implement training infrastructure spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training for Physical AI world foundation models
  • Develop and improve pre-training and SFT pipelines, including large-scale data loading, distributed training, and checkpointing
  • Develop and improve the inference and evaluation stack, including inference engines, inference/generation pipelines, RL rollout support, and evaluation pipelines
  • Use continuous batching and KV-cache management to achieve high throughput and low latency
  • Build and improve interaction and data flow among RL system roles: policy, rollout, reward, and simulation
  • Integrate and orchestrate simulation and robotics environments as RL environments, driving the simulation–rollout–training loop at scale
  • Build and refine the distributed training backend, including sharding/parallelism, mixed precision, activation checkpointing, and multi-GPU memory/throughput optimization
  • Improve efficiency, scalability, and resiliency of training and RL workloads through fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure
  • Define actionable reliability and efficiency metrics
  • Root-cause, triage, and resolve failures from the application level through framework, GPU, network, and hardware levels

Requirements

What you’ll need
  • 5+ years developing software infrastructure for large-scale AI or distributed systems
  • Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience
  • Strong debugging and triage skills across the stack, from AI application to GPU/hardware behavior
  • Proven track record building and scaling large-scale distributed systems, ideally distributed training or inference
  • Hands-on experience with AI training and/or inference infrastructure, RL/post-training, training frameworks, or inference serving
  • Proficiency in Python and scripting
  • Solid software engineering practices including testing, defensive programming, version control, and CI
  • Experience building RL/post-training infrastructure, including PPO/GRPO/DPO pipelines, rollout engines, and asynchronous RL
  • Experience with production-grade pre-training/SFT infrastructure
  • Experience integrating simulation/robotics environments into training or RL loops, including vectorized environments and sim-to-real workflows
  • Knowledge of deep learning framework internals, PyTorch (FSDP/DTensor), Megatron or equivalent, distributed training, and related optimization techniques
  • Proficiency in C/C++/CUDA for performance-critical components and custom kernels

Benefits

Comp & perks
  • Highly competitive salaries
  • Comprehensive benefits package
  • Benefits for you and your family