Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Performance Engineer – DGX Cloud

NVIDIA

Senior Performance Engineer analyzing performance of AI workloads across various stacks at NVIDIA. Collaborating with engineers and architects to achieve optimization in DGX Cloud systems.

Posted 7/28/2026full-timeSanta Clara • California, Oregon, Texas, Washington • 🇺🇸 United StatesSenior💰 $224,000 - $431,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Expertise in performance engineering and optimization of large-scale AI workloads, with strong programming skills in C++ and Python. Proficient in using deep learning frameworks and GPU computing systems to analyze and enhance system performance.

Highest-signal resume keywords
C++ ProgrammingPython ProgrammingPerformance EngineeringCUDADeep Learning Frameworks

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Performance AnalysisBenchmarkingProfilingOptimizationDistributed SystemsOperating SystemsWorkload CharacterizationData AnalysisAutomation WorkflowsSystem-Level Performance Analysis
Soft Skills
CommunicationCollaborationPrioritizationInfluencing
Tools & Technologies
PyTorchJAX/XLAGPU Computing Systems
Industry Keywords
AI WorkloadsLarge-Scale ClustersDeep LearningPerformance Metrics

Tech Stack

Tools & technologies
Distributed SystemsPythonPyTorch

About the role

Key responsibilities & impact
  • Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks.
  • Design and execute rigorous performance studies to establish baselines, diagnose regressions, and quantify bottlenecks.
  • Define performance and efficiency evaluation methodologies, benchmarks, and success metrics for AI workloads.
  • Use profiling, observability, and data analysis to turn performance measurements into actionable optimization plans.
  • Partner with deep learning engineers, platform teams, and GPU architects to validate and deliver performance improvements.
  • Communicate performance findings, trade-offs, and recommendations clearly to influence system and software design decisions.

Requirements

What you’ll need
  • BS or higher degree in computer science, computer engineering, or a related field (or equivalent experience).
  • 12+ years of experience in strong programming skills in C++ and Python, with the ability to build reliable analysis and automation workflows
  • Solid foundation in operating systems, computer architecture, and distributed systems
  • Experience with performance engineering, benchmarking, profiling, and optimization of complex software or systems
  • Ability to communicate technical findings, prioritize high-impact work, and build alignment across teams
  • Experience analyzing large-scale AI clusters or distributed training and inference workloads (Ways to stand out from the crowd).
  • Experience with CUDA, GPU computing systems, and GPU performance analysis.
  • Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA.
  • Deep understanding of system-level performance analysis, workload characterization, and optimization.

Benefits

Comp & perks
  • Eligible for equity and benefits