FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Solutions Architect – Large Scale AI Training
NVIDIASenior Solutions Architect guiding EMEA AI model builders on distributed training, alignment, and NVIDIA’s full training stack. Optimizing large-scale foundation-model infrastructure and shaping product roadmaps.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in distributed AI training, particularly with frameworks like Megatron-LM and NeMo, while effectively optimizing large-scale training workflows and infrastructure. Strong communication skills facilitate collaboration with research scientists and engineers to align product development with customer needs.
Highest-signal resume keywords
Distributed AI TrainingMegatron-LMNeMoReinforcement LearningHPC Environments
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed Training FrameworksTraining Infrastructure OptimizationGPU UtilizationMemory ManagementFine-Tuning with Reinforcement LearningExpert Load BalancingSpeculative DecodingMulti-Node GPU ClustersLarge-Scale Training WorkflowsPost-Training Recipes
Soft Skills
Excellent Communication SkillsCollaboration
Tools & Technologies
PyTorchRL/Gym
Certifications & Qualifications
MS or PhD in Computer ScienceEngineering
Industry Keywords
AI Model BuildersAI DatacenterHPC CenterFrontier AI LabPublished WorkOpen-Source Contributions
Tech Stack
Tools & technologiesNode.jsPyTorch
About the role
Key responsibilities & impact- Build and manage strategic technical relationships with leading EMEA AI model builders developing large-scale foundation models
- Define software stacks and infrastructure for large-scale training and post-training workflows, including Reinforcement Learning
- Guide customers on distributed training strategies and efficient large-scale training and post-training recipes using PyTorch, Megatron-LM, or NeMo (RL/Gym)
- Help customers optimize training and fine-tuning efficiency at scale, including GPU utilization, communication overlap, and memory management
- Collaborate with NVIDIA product and research teams to present customer needs
- Craft NVIDIA product roadmap based on customer feedback
- Build or support hackathons, demos, and technical conferences to animate the developer community
Requirements
What you’ll need- MS or PhD in Computer Science, Engineering, or equivalent experience
- Over 7 years of practical experience in distributed AI training
- Direct involvement with HPC and/or AI environments with multi-node GPU clusters
- Solid understanding of training infrastructure and its effects on efficiency and scalability
- Strong proficiency with Megatron-LM, NeMo, or equivalent distributed training frameworks
- Excellent communication skills with ability to engage research scientists and infrastructure engineers
- Experience in fine-tuning with Reinforcement Learning (RLVR, RLHF) at scale
- Experience with LatentMoE, expert load balancing, and speculative decoding for MoE inference
- Prior experience in an AI Datacenter/HPC center, national lab, or frontier AI lab environment
- Published work or open-source contributions in distributed training
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package
- Equal opportunity employer