Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Principal Software Engineer – Compute Infrastructure

NVIDIA

Principal Software Engineer architecting NVIDIA’s global GPU and AI compute infrastructure. Leading Kubernetes platforms, frontier AI inference operations, capacity planning, and large-scale workload migrations.

Posted 8/31/2026full-timeSanta Clara • California • 🇺🇸 United StatesLead💰 $248,000 - $391,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Kubernetes architecture, automation at scale, and advanced compute platform engineering, with a strong focus on operational maturity and self-service architectures. Proven ability to lead technical direction and manage complex global environments with a deep understanding of hardware technologies and virtualization.

Highest-signal resume keywords
Kubernetes ArchitectureInfrastructure-As-Code DevelopmentCompute Platform EngineeringAutomation At ScaleLeadership In Technical Direction

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesOpenShiftKubeVirtTerraformGoPythonGitOpsNFSv4NVMe/TCPHyperconverged Storage
Soft Skills
LeadershipInfluencing Technical Direction
Tools & Technologies
ArgoCDCloud PlatformsVirtualization InfrastructureTelemetry SystemsAutomated Remediation Pipelines
Industry Keywords
AI Inference SystemsCapacity PlanningService Level AgreementsMulti-Cloud DeploymentLegacy Workload Migration

Tech Stack

Tools & technologies
AWSCloudGoGoogle Cloud PlatformKubernetesMicroservicesOpenShiftPythonTerraform

About the role

Key responsibilities & impact
  • Lead the architectural vision for a massive global platform and spearhead operationalization of internal frontier-class AI inference systems
  • Architect and transform the global enterprise compute platform running thousands of nodes and tens of thousands of VMs and containers via OpenShift and KubeVirt
  • Define service tiers, SLAs, and automated cluster lifecycles
  • Build automated remediation pipelines, hardware watchdogs, and telemetry for pre-release, rack-scale GPU systems
  • Collect and review system data for capacity planning amid extreme hardware supply constraints
  • Develop strategies including public cloud bursting, hardware dogfooding, and evaluation of alternative compute architectures such as ARM
  • Drive adoption of standard platforms through self-service architectures, APIs, and Terraform/OpenTofu providers
  • Evaluate application architectures and lead migrations of massive legacy workloads, including large-scale long-running VDI environments, into modern Kubernetes orchestration

Requirements

What you’ll need
  • Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience
  • 15+ years of proven experience in compute platform engineering, site reliability, or systems architecture with a heavy focus on automation at massive scale
  • Deep expertise in Kubernetes architecture and designing/deploying virtualization architectures, specifically operating VMs inside K8s (KubeVirt, OpenShift)
  • In-depth knowledge of hardware technologies (GPUs, high-speed backplane networking) with a track record of mitigating hardware-level failures, silent data corruption, and anomalies in large-scale environments
  • Experience running large global environments spanning bare metal, virtualized infrastructure, and cloud with a unified GitOps posture (ArgoCD or similar)
  • Proficiency in programming languages such as Go and/or Python, alongside expert-level infrastructure-as-code development (Terraform, Config Management)
  • Strong leadership skills with the ability to influence technical direction across highly autonomous teams without relying on top-down mandates
  • Hands-on experience managing bleeding-edge, pre-release hardware in production environments
  • Deep understanding of advanced storage migrations and protocols (NFSv4, NVMe/TCP, Hyperconverged storage)
  • Solid understanding of microservices architecture and seamless multi-cloud deployment strategies (AWS, GCP)
  • Proven track record of building “Day 2” operational maturity (self-service, advanced auto-remediation, strict SLAs) from the ground up on existing foundations

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score