Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Solutions Architect, Cloud Partner Operations

NVIDIA

NVIDIA Solutions Architect improving Day 2 operations for AI cloud partners. Advancing reliability, performance, economics, and operational maturity across large-scale AI infrastructure.

Posted 8/13/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $224,000 - $356,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates deep expertise in Day 2 operations for large-scale GPU and cloud infrastructure, with a focus on improving reliability, performance, and operational practices. Proven ability to lead cross-functional teams and drive adoption of new technologies in production environments.

Highest-signal resume keywords
Large-Scale GPU Infrastructure ExperienceCloud Engineering ExpertiseKubernetes and Slurm ProficiencyAutomation with Terraform and AnsibleStrong Linux Knowledge

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production InfrastructureSite Reliability EngineeringDistributed InfrastructurePython ScriptingBash ScriptingEvidence-Led TroubleshootingCapacity PlanningPerformance OptimizationIncident ResponseAutomation
Soft Skills
Strong CommunicationPrioritization SkillsTime-Management SkillsCollaborationLeadership
Tools & Technologies
DCGMInfiniBandNCCLPrometheusGrafanaOpenTelemetryTerraformAnsibleArgo CDIBM Storage Scale
Industry Keywords
HPCAI InfrastructureNVIDIA PlatformsCloud Partner OperationsFleet HealthUnit Economics24/7 OperationsHigh-Speed EthernetLustreVAST Data

Tech Stack

Tools & technologies
AnsibleCloudGrafanaKubernetesLinuxPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Solve hard Day 2 operations problems at scale alongside partner engineers.
  • Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices.
  • Prepare partners for new NVIDIA platforms, capacity, services, and use cases.
  • Drive adoption in live environments without degrading service.
  • Improve reliability, performance, utilization, recovery time, and cost per token.
  • Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response.
  • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows.
  • Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering.
  • Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.

Requirements

What you’ll need
  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.
  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
  • Experience building, operating, or improving distributed infrastructure under real production load.
  • Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience.
  • Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
  • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling.
  • Strong Linux knowledge.
  • Experience with Python, Bash, or similar scripting for automation.
  • Evidence-led troubleshooting across system boundaries.
  • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority.
  • Strong communication, prioritization, and time-management skills across multiple partner engagements.
  • Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.

Benefits

Comp & perks
  • Competitive salaries
  • Generous benefits package
  • Equity