Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA

Distinguished Engineer defining cluster operations and production standards for NVIDIA’s DGX Cloud GPU infrastructure. Leading architecture, automation, reliability, and cross-team execution at scale.

Posted 9/3/2026full-timeSanta Clara • California, Oregon • 🇺🇸 United StatesSeniorLead💰 $320,000 - $488,750 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in defining technical strategies for large-scale distributed systems and cloud platforms, with a focus on operational excellence, architectural direction, and automation. Proven ability to lead cross-organizational technical efforts and improve production readiness through effective collaboration and engineering standards.

Highest-signal resume keywords
Kubernetes-Based Production SystemsInfrastructure AutomationDistributed Systems OperationsSoftware Engineering in Python or GoTechnical Leadership in Production Engineering

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed SystemsInfrastructure PlatformsProduction EnvironmentsOperational WorkflowsAPIsService InterfacesAutomation FrameworksLinuxNetworkingProduction Reliability
Soft Skills
Cross-Team CollaborationTechnical JudgmentProblem-Solving
Tools & Technologies
KubernetesCloud PlatformsAutomation Tools
Industry Keywords
Technical StrategyArchitectural DirectionOperational StandardsProduction ReadinessService Reliability

Tech Stack

Tools & technologies
CloudDistributed SystemsGoKubernetesLinuxPython

About the role

Key responsibilities & impact
  • Define the long-range technical strategy for operating DGX Cloud clusters consistently across local data centers, hyperscalers, and NeoCloud environments
  • Define architectural direction and operating standards for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability
  • Guide roadmaps and execution of cross-organizational investments improving production readiness, operational safety, performance, and coordination
  • Make and influence high-impact technical decisions across platform, hardware, provider, and service teams
  • Build workflows, interfaces, and engineering handshakes across Kubernetes production service, provider and hardware preparation, on-prem, and bare-metal operations
  • Restructure operations and service-layer reliability domains
  • Build and evolve automation, APIs, operating workflows, and readiness gates for moving new capacity into stable production and balancing existing capacity
  • Implement production operating approaches that reduce manual input and increase ownership clarity, consistency, traceability, and release safety
  • Partner with platform, hardware, provider engineering, service owners, and Production Engineering leaders to convert recurring friction into durable improvements
  • Raise standards for operability, resilience, scalability, and performance through build leadership, architecture review, and technical standards

Requirements

What you’ll need
  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments
  • Company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms
  • Track record of defining operating models, architectural direction, and engineering standards across multiple technical domains and organizations
  • Record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating for clarity, and delivering measurable outcomes
  • Deep experience with Kubernetes-based production systems, infrastructure automation, or distributed systems operations
  • Strong software engineering skills in Python, Go, or similar low-level programming languages
  • Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns
  • Experience crafting operational workflows, APIs, service interfaces, or automation frameworks that become the standard way teams run production systems
  • Strong architectural judgment and a validated history of simplifying complex operational problems through reusable software, clear technical strategy, and durable engineering direction

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score