FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Research Scientist, Synthetic Data Generation
NVIDIAPrincipal scientist leading NVIDIA’s synthetic data generation for frontier LLM training. Building open-source NeMo tools across text, code, multimodal, agentic, and privacy-preserving datasets.
Posted 9/4/2026full-timeRemote • California, Colorado, Massachusetts, New York • 🇺🇸 United StatesLead💰 $272,000 - $431,250 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in synthetic data generation, multimodal machine learning, and large language models (LLMs), with a strong focus on developing open-source libraries and scalable data pipelines. Proven ability to mentor teams and publish research in leading AI and machine learning conferences.
Highest-signal resume keywords
PhD In Computer Science15+ Years Of Engineering ExperienceSynthetic Data GenerationMultimodal Machine LearningDifferential Privacy
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Synthetic Data GenerationGenerative ModelingLarge Language Models (LLMs)Data Pipeline OptimizationReinforcement Learning (RL)Open-Source Library DevelopmentAutomated Quality EvaluationMultimodal Data GenerationFunction CallingReward Modeling
Soft Skills
MentoringCollaboration
Tools & Technologies
NVIDIA NeMoGitCI-CDVLLMTGI
Industry Keywords
Machine LearningStatisticsHealthcareFinanceGovernment
About the role
Key responsibilities & impact- Set the technical direction for synthetic data generation across NVIDIA's frontier model efforts
- Define and build open-source libraries within the NVIDIA NeMo ecosystem
- Generate synthetic datasets across text, code, structured, and multimodal data for LLM pre- and post-training
- Build and scale LLM-based data generation pipelines with automated quality evaluation
- Develop synthetic trajectories, multi-turn interactions, function calling, executable environments, reward modeling, and verifiable-reward data for agentic and tool-use training
- Advance multimodal synthetic data generation for image, document, video, and audio
- Advance privacy-preserving and safe synthesis using differential privacy, anonymization, and de-identification
- Develop and maintain open-source libraries and SDKs with clean APIs and strong documentation
- Drive software excellence through modern tooling, configurable architecture, and professional Git/CI-CD
- Publish original research at leading machine learning and AI conferences
- Collaborate with research, engineering, product, model teams, and external labs
- Mentor scientists and engineers across the team
Requirements
What you’ll need- PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience
- 15+ years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related areas
- Deep technical understanding of LLMs, pre-training, post-training, RL stages, and inference frameworks such as vLLM or TGI
- Proven track record of developing or maintaining software libraries used by a broad developer community
- Experience building and optimizing scalable data pipelines for large-scale model training, including throughput, distributed inference, and cost at cluster scale
- Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL or similar
- Significant open-source contributions in ML or data tooling, with community adoption
- Experience with multimodal generation or understanding, including vision-language, document AI, video, or audio
- Experience generating data for agentic, tool-use, or reinforcement-learning post-training, including RL environment design
- Background in differential privacy, de-identification, or synthetic data for regulated industries such as healthcare, finance, or government
- Experience influencing model training decisions at frontier scale, or partnering directly with pre-training and post-training teams
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score