FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Infrastructure Software Engineer
NVIDIAAI infrastructure engineer building scalable pre-training, inference, and reinforcement-learning systems for NVIDIA’s Physical AI foundation models. Optimizing distributed GPU training, simulation integration, reliability, and performance.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and implementing training infrastructure for AI models, with a strong focus on reinforcement learning and distributed systems. Proficient in Python, C/C++/CUDA, and deep learning frameworks, with a proven ability to optimize performance and scalability in large-scale environments.
Highest-signal resume keywords
Python ProficiencyDistributed Systems DevelopmentReinforcement Learning InfrastructureDeep Learning Frameworks (PyTorch)Debugging and Triage Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software Infrastructure DevelopmentLarge-Scale Data LoadingCheckpointingInference Engine DevelopmentMixed Precision OptimizationAsynchronous Reinforcement LearningProduction-Grade InfrastructureTesting and Defensive ProgrammingVersion ControlCI Practices
Tools & Technologies
PyTorch (FSDP/DTensor)MegatronCUDASimulation EnvironmentsVectorized Environments
Certifications & Qualifications
Bachelor's Degree in Computer Science or Related Field
Industry Keywords
AI TrainingDistributed TrainingInference ServingSimulation-Robotics IntegrationReliability Metrics
Tech Stack
Tools & technologiesDistributed SystemsPythonPyTorch
About the role
Key responsibilities & impact- Create and implement training infrastructure spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training for Physical AI world foundation models
- Develop and improve pre-training and SFT pipelines, including large-scale data loading, distributed training, and checkpointing
- Develop and improve the inference and evaluation stack, including inference engines, inference/generation pipelines, RL rollout support, and evaluation pipelines
- Use continuous batching and KV-cache management to achieve high throughput and low latency
- Build and improve interaction and data flow among RL system roles: policy, rollout, reward, and simulation
- Integrate and orchestrate simulation and robotics environments as RL environments, driving the simulation–rollout–training loop at scale
- Build and refine the distributed training backend, including sharding/parallelism, mixed precision, activation checkpointing, and multi-GPU memory/throughput optimization
- Improve efficiency, scalability, and resiliency of training and RL workloads through fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure
- Define actionable reliability and efficiency metrics
- Root-cause, triage, and resolve failures from the application level through framework, GPU, network, and hardware levels
Requirements
What you’ll need- 5+ years developing software infrastructure for large-scale AI or distributed systems
- Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience
- Strong debugging and triage skills across the stack, from AI application to GPU/hardware behavior
- Proven track record building and scaling large-scale distributed systems, ideally distributed training or inference
- Hands-on experience with AI training and/or inference infrastructure, RL/post-training, training frameworks, or inference serving
- Proficiency in Python and scripting
- Solid software engineering practices including testing, defensive programming, version control, and CI
- Experience building RL/post-training infrastructure, including PPO/GRPO/DPO pipelines, rollout engines, and asynchronous RL
- Experience with production-grade pre-training/SFT infrastructure
- Experience integrating simulation/robotics environments into training or RL loops, including vectorized environments and sim-to-real workflows
- Knowledge of deep learning framework internals, PyTorch (FSDP/DTensor), Megatron or equivalent, distributed training, and related optimization techniques
- Proficiency in C/C++/CUDA for performance-critical components and custom kernels
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package
- Benefits for you and your family