Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Solutions Architect, Cloud Partner Operations

NVIDIA

NVIDIA Solutions Architect improving Day 2 operations across AI cloud partners. Scaling reliability, performance, efficiency, and operational maturity for GPU infrastructure.

Posted 8/21/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $224,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating large-scale GPU and AI infrastructure, with a focus on Day 2 operations, incident management, and automation. Proven ability to lead cross-functional teams and improve operational practices while ensuring reliability and performance.

Highest-signal resume keywords
Large-Scale GPU InfrastructureKubernetes or SlurmIncident ManagementAutomation with Terraform or AnsibleLinux and Python Proficiency

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production InfrastructureCloud EngineeringSolutions ArchitectureSite Reliability EngineeringHPCDistributed InfrastructureTroubleshootingBenchmarkingInfrastructure as CodeAutomated Diagnosis
Soft Skills
Strong CommunicationPrioritizationTime ManagementCollaborationLeadership
Tools & Technologies
PrometheusGrafanaOpenTelemetryTerraformAnsibleArgo CDNVIDIA Rack-Scale PlatformsSpectrum-XUFMBase Command Manager
Industry Keywords
Day 2 OperationsOperational PracticesGPU SchedulingMulti-TenancyObservabilityIncident Response24/7 Operations FunctionFleet HealthUnit EconomicsAgent-Based Remediation

Tech Stack

Tools & technologies
AnsibleCloudGrafanaKubernetesLinuxPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Solve hard Day 2 operations problems at scale alongside partner engineers
  • Find causes, prototype approaches, validate solutions under representative load, and leave operational practices partners can run
  • Help partners prepare operating models for new NVIDIA platforms, capacity, services, and use cases
  • Drive adoption in live environments without degrading service
  • Improve reliability, performance, and economics using incident frequency, recovery time, utilization, and cost-per-token measures
  • Identify and help close Day 2 maturity gaps across people, process, tooling, telemetry, security, and incident response
  • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows
  • Spot cross-partner patterns and provide field evidence to account teams, support, product, and engineering
  • Improve NVIDIA's factory planning function

Requirements

What you’ll need
  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience
  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure
  • Experience building, operating, or improving distributed infrastructure under real production load
  • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure
  • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling
  • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation
  • Detailed evidence-led troubleshooting across system boundaries
  • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority
  • Strong communication, prioritization, and time-management skills across multiple partner engagements
  • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load
  • Experience building or maturing a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design
  • Hands-on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72, or NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and GPU or Network Operators
  • Experience improving fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation

Benefits

Comp & perks
  • Competitive salaries
  • Generous benefits package
  • Equity
  • Benefits