FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Solutions Architect, Cloud Partner Operations
NVIDIANVIDIA Solutions Architect improving Day 2 operations across AI cloud partners. Scaling reliability, performance, efficiency, and operational maturity for GPU infrastructure.
Posted 8/21/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $224,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating large-scale GPU and AI infrastructure, with a focus on Day 2 operations, incident management, and automation. Proven ability to lead cross-functional teams and improve operational practices while ensuring reliability and performance.
Highest-signal resume keywords
Large-Scale GPU InfrastructureKubernetes or SlurmIncident ManagementAutomation with Terraform or AnsibleLinux and Python Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production InfrastructureCloud EngineeringSolutions ArchitectureSite Reliability EngineeringHPCDistributed InfrastructureTroubleshootingBenchmarkingInfrastructure as CodeAutomated Diagnosis
Soft Skills
Strong CommunicationPrioritizationTime ManagementCollaborationLeadership
Tools & Technologies
PrometheusGrafanaOpenTelemetryTerraformAnsibleArgo CDNVIDIA Rack-Scale PlatformsSpectrum-XUFMBase Command Manager
Industry Keywords
Day 2 OperationsOperational PracticesGPU SchedulingMulti-TenancyObservabilityIncident Response24/7 Operations FunctionFleet HealthUnit EconomicsAgent-Based Remediation
Tech Stack
Tools & technologiesAnsibleCloudGrafanaKubernetesLinuxPrometheusPythonTerraform
About the role
Key responsibilities & impact- Solve hard Day 2 operations problems at scale alongside partner engineers
- Find causes, prototype approaches, validate solutions under representative load, and leave operational practices partners can run
- Help partners prepare operating models for new NVIDIA platforms, capacity, services, and use cases
- Drive adoption in live environments without degrading service
- Improve reliability, performance, and economics using incident frequency, recovery time, utilization, and cost-per-token measures
- Identify and help close Day 2 maturity gaps across people, process, tooling, telemetry, security, and incident response
- Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows
- Spot cross-partner patterns and provide field evidence to account teams, support, product, and engineering
- Improve NVIDIA's factory planning function
Requirements
What you’ll need- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience
- 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure
- Experience building, operating, or improving distributed infrastructure under real production load
- Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure
- Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling
- Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation
- Detailed evidence-led troubleshooting across system boundaries
- Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority
- Strong communication, prioritization, and time-management skills across multiple partner engagements
- Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load
- Experience building or maturing a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design
- Hands-on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72, or NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and GPU or Network Operators
- Experience improving fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation
Benefits
Comp & perks- Competitive salaries
- Generous benefits package
- Equity
- Benefits