FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Software Engineer – Compute Infrastructure
NVIDIAPrincipal Software Engineer architecting NVIDIA’s global GPU and AI compute infrastructure. Leading Kubernetes platforms, frontier AI inference operations, capacity planning, and large-scale workload migrations.
Posted 8/31/2026full-timeSanta Clara • California • 🇺🇸 United StatesLead💰 $248,000 - $391,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Kubernetes architecture, automation at scale, and advanced compute platform engineering, with a strong focus on operational maturity and self-service architectures. Proven ability to lead technical direction and manage complex global environments with a deep understanding of hardware technologies and virtualization.
Highest-signal resume keywords
Kubernetes ArchitectureInfrastructure-As-Code DevelopmentCompute Platform EngineeringAutomation At ScaleLeadership In Technical Direction
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesOpenShiftKubeVirtTerraformGoPythonGitOpsNFSv4NVMe/TCPHyperconverged Storage
Soft Skills
LeadershipInfluencing Technical Direction
Tools & Technologies
ArgoCDCloud PlatformsVirtualization InfrastructureTelemetry SystemsAutomated Remediation Pipelines
Industry Keywords
AI Inference SystemsCapacity PlanningService Level AgreementsMulti-Cloud DeploymentLegacy Workload Migration
Tech Stack
Tools & technologiesAWSCloudGoGoogle Cloud PlatformKubernetesMicroservicesOpenShiftPythonTerraform
About the role
Key responsibilities & impact- Lead the architectural vision for a massive global platform and spearhead operationalization of internal frontier-class AI inference systems
- Architect and transform the global enterprise compute platform running thousands of nodes and tens of thousands of VMs and containers via OpenShift and KubeVirt
- Define service tiers, SLAs, and automated cluster lifecycles
- Build automated remediation pipelines, hardware watchdogs, and telemetry for pre-release, rack-scale GPU systems
- Collect and review system data for capacity planning amid extreme hardware supply constraints
- Develop strategies including public cloud bursting, hardware dogfooding, and evaluation of alternative compute architectures such as ARM
- Drive adoption of standard platforms through self-service architectures, APIs, and Terraform/OpenTofu providers
- Evaluate application architectures and lead migrations of massive legacy workloads, including large-scale long-running VDI environments, into modern Kubernetes orchestration
Requirements
What you’ll need- Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience
- 15+ years of proven experience in compute platform engineering, site reliability, or systems architecture with a heavy focus on automation at massive scale
- Deep expertise in Kubernetes architecture and designing/deploying virtualization architectures, specifically operating VMs inside K8s (KubeVirt, OpenShift)
- In-depth knowledge of hardware technologies (GPUs, high-speed backplane networking) with a track record of mitigating hardware-level failures, silent data corruption, and anomalies in large-scale environments
- Experience running large global environments spanning bare metal, virtualized infrastructure, and cloud with a unified GitOps posture (ArgoCD or similar)
- Proficiency in programming languages such as Go and/or Python, alongside expert-level infrastructure-as-code development (Terraform, Config Management)
- Strong leadership skills with the ability to influence technical direction across highly autonomous teams without relying on top-down mandates
- Hands-on experience managing bleeding-edge, pre-release hardware in production environments
- Deep understanding of advanced storage migrations and protocols (NFSv4, NVMe/TCP, Hyperconverged storage)
- Solid understanding of microservices architecture and seamless multi-cloud deployment strategies (AWS, GCP)
- Proven track record of building “Day 2” operational maturity (self-service, advanced auto-remediation, strict SLAs) from the ground up on existing foundations
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score