Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA

Distinguished Engineer defining production architecture and automation for NVIDIA’s DGX Cloud GPU infrastructure. Leading cross-team reliability, lifecycle, and operational standards at scale.

Posted 9/2/2026full-timeRemote • California, New York, South Dakota, Wyoming • 🇺🇸 United StatesSeniorLead💰 $320,000 - $488,750 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in defining architectural direction and operational guidelines for large-scale distributed systems, particularly in DGX Cloud environments. Proven ability to lead cross-organizational technical efforts, ensuring production readiness and operational excellence.

Highest-signal resume keywords
Technical LeadershipKubernetes Service ManagementLarge-Scale Distributed SystemsInfrastructure OperationsAutomation and Workflow Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringSystem KnowledgeProduction InsightArchitectural DirectionOperating ModelsEngineering StandardsAPIs DevelopmentInfrastructure PlatformsProduction EnvironmentsCloud Platforms
Soft Skills
Cross-Team CoordinationInfluential LeadershipComplex Problem SolvingCollaboration
Tools & Technologies
DGX CloudOn-Premises InfrastructureHyperscaler EnvironmentsNeoCloudBare-Metal Infrastructure
Certifications & Qualifications
BS in Computer ScienceMS in Computer SciencePhD in Computer ScienceEquivalent Experience
Industry Keywords
Production EngineeringSite Reliability Engineering (SRE)Infrastructure SoftwareCloud Platforms

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetes

About the role

Key responsibilities & impact
  • Define the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments
  • Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources
  • Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination
  • Make and guide high-impact technical decisions resolving how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production
  • Develop robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains
  • Act as a senior technical leader in the Production Engineering group
  • Build the architectural direction for cluster operations in DGX Cloud
  • Set operating standards and direct the evolution of the production model
  • Drive delivery of cross-organizational capabilities ensuring DGX Cloud resources remain usable, maintainable, and continuously improved at scale
  • Lead by influence across several teams and critical production results

Requirements

What you’ll need
  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments
  • Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms
  • Confirmed experience in establishing operating models, architectural direction, and engineering standards across various technical domains and organizations
  • Consistent record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating complexity and delivering measurable outcomes
  • Deep software engineering expertise, system knowledge, and production insight
  • Experience with Kubernetes service management
  • Experience with on-premises, hyperscaler, NeoCloud, and bare-metal infrastructure operations
  • Experience developing automation, workflows, interfaces, APIs, architectures, or operating standards for large-scale infrastructure environments

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score