FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer, Evaluation Flywheel – Autonomous Vehicles
NVIDIASenior Software Engineer responsible for the eval flywheel architecture and quality in NVIDIA's self-driving technology. Collaborating with teams to ensure trust and high standards in evaluation processes.
Posted 7/28/2026full-timeSanta Clara • California, District of Columbia, North Carolina, Washington • 🇺🇸 United StatesSenior💰 $224,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and evaluating machine learning systems, with a focus on golden dataset curation, metric performance measurement, and data pipeline development. Strong collaboration with cross-functional teams to ensure evaluation quality and system reliability.
Highest-signal resume keywords
Machine Learning EvaluationGolden Dataset CurationMetric Performance MeasurementPython ProgrammingData Engineering
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine LearningMetric DesignGround-Truth CurationPrecision/Recall MethodologyData Pipeline DevelopmentSoftware DevelopmentRoboticsAutonomous VehiclesLarge-Scale ML SystemsQuality Reporting
Soft Skills
CollaborationCommunicationStrategic ThinkingProblem Solving
Tools & Technologies
Data PipelinesSelf-Serve Dataset ToolsQuality Reporting Tools
Industry Keywords
Autonomous VehiclesRoboticsMachine Learning SystemsEvaluation QualityVersioningRelease Processes
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Owning the eval flywheel's strategy and architecture: how road and simulation driving data becomes curated golden datasets, how metrics are measured against them (precision/recall), and how those results earn lasting trust with the teams that depend on them.
- Setting the standard for evaluation quality: golden dataset curation, versioning, and health; metric performance measurement; and release processes that keep results dependable as the system evolves.
- Building the tooling that helps our metric developers iterate quickly: self-serve dataset pipelines, metric performance measurement, and quality reporting used every day by the team and our partners.
- Partnering with senior engineers and leaders across test engineering, behavior planning, and infrastructure — setting expectations, working through trade-offs, and being the voice of evaluation quality in cross-team decisions.
- Working directly with AI model developers so evaluation iteration speed becomes an advantage for the whole program, including our push into learned, VLM-based evaluation.
Requirements
What you’ll need- 12+ years building software, with significant time in autonomous vehicles, robotics, or large-scale ML systems.
- Deep experience evaluating ML or robotic systems: metric design, ground-truth and golden dataset curation, precision/recall methodology, and the data pipelines behind them.
- Strong Python and data engineering skills for production-scale pipelines.
- BS or MS in Computer Science, Robotics, or a related field (or equivalent experience).
Benefits
Comp & perks- equity
- benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score