FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI and ML Infra Software Engineer, GPU Clusters
NVIDIAAI Infrastructure Software Engineer collaborating with researchers to enhance AI/ML infrastructure at NVIDIA. Focused on optimizing performance and facilitating innovative AI research on GPU clusters.
Posted 7/23/2026full-timeSanta Clara • California, Washington • 🇺🇸 United StatesMid-LevelSenior💰 $124,000 - $195,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI/ML infrastructure, including High Performance Computing (HPC) and accelerated computing technologies. Proficient in optimizing large-scale distributed training workloads and collaborating with diverse teams to enhance AI researcher efficiency.
Highest-signal resume keywords
AI/ML InfrastructureHigh Performance Computing (HPC)Distributed Training Workloads OptimizationProgramming Languages (Python, Go, Bash)Cloud Computing Platforms (AWS, GCP, Azure)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI/ML WorkloadsAccelerated ComputingStorage Technologies (Lustre, GPFS, BeeGFS)Scheduling & Orchestration (Slurm, Kubernetes, LSF)High-Speed Networking (Infiniband, RoCE, Amazon EFA)Container Technologies (Docker, Enroot)Parallel Computing FrameworksData ProcessingModel TrainingInference Pipelines
Soft Skills
Excellent CommunicationCollaboration SkillsPassion for Learning
Certifications & Qualifications
MS or PhD in Computer Science or Related Field
Tech Stack
Tools & technologiesAWSAzureCloudDockerGoGoogle Cloud PlatformKubernetesPythonPyTorch
About the role
Key responsibilities & impact- Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles
- Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization
- Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results
- Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem
- Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company
Requirements
What you’ll need- Recent graduate with a MS, PhD or equivalent experience in Computer Science or related field
- Proven experience in AI/ML and HPC workloads and infrastructure
- Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure
- In-depth knowledge of accelerated computing (e.g., GPU, custom silicon)
- Storage (e.g., Lustre, GPFS, BeeGFS)
- Scheduling & orchestration (e.g., Slurm, Kubernetes, LSF)
- High-speed networking (e.g., Infiniband, RoCE, Amazon EFA)
- Containers technologies (Docker, Enroot)
- Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX
- Deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines
- Proficiency in programming & scripting languages such as Python, Go, Bash
- Familiarity with cloud computing platforms (e.g., AWS, GCP, Azure)
- Experience with parallel computing frameworks and paradigms.
- Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field.
- Excellent communication and collaboration skills
Benefits
Comp & perks- competitive salaries
- comprehensive benefits package