FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Solutions Architect, Cloud Partner Operations
NVIDIANVIDIA Solutions Architect improving Day 2 operations for AI cloud partners. Advancing reliability, performance, economics, and operational maturity across large-scale AI infrastructure.
Posted 8/13/2026full-timeRemote • California • 🇺🇸 United StatesSenior💰 $224,000 - $356,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates deep expertise in Day 2 operations for large-scale GPU and cloud infrastructure, with a focus on improving reliability, performance, and operational practices. Proven ability to lead cross-functional teams and drive adoption of new technologies in production environments.
Highest-signal resume keywords
Large-Scale GPU Infrastructure ExperienceCloud Engineering ExpertiseKubernetes and Slurm ProficiencyAutomation with Terraform and AnsibleStrong Linux Knowledge
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production InfrastructureSite Reliability EngineeringDistributed InfrastructurePython ScriptingBash ScriptingEvidence-Led TroubleshootingCapacity PlanningPerformance OptimizationIncident ResponseAutomation
Soft Skills
Strong CommunicationPrioritization SkillsTime-Management SkillsCollaborationLeadership
Tools & Technologies
DCGMInfiniBandNCCLPrometheusGrafanaOpenTelemetryTerraformAnsibleArgo CDIBM Storage Scale
Industry Keywords
HPCAI InfrastructureNVIDIA PlatformsCloud Partner OperationsFleet HealthUnit Economics24/7 OperationsHigh-Speed EthernetLustreVAST Data
Tech Stack
Tools & technologiesAnsibleCloudGrafanaKubernetesLinuxPrometheusPythonTerraform
About the role
Key responsibilities & impact- Solve hard Day 2 operations problems at scale alongside partner engineers.
- Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices.
- Prepare partners for new NVIDIA platforms, capacity, services, and use cases.
- Drive adoption in live environments without degrading service.
- Improve reliability, performance, utilization, recovery time, and cost per token.
- Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response.
- Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows.
- Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering.
- Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.
Requirements
What you’ll need- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.
- 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
- Experience building, operating, or improving distributed infrastructure under real production load.
- Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience.
- Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
- Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling.
- Strong Linux knowledge.
- Experience with Python, Bash, or similar scripting for automation.
- Evidence-led troubleshooting across system boundaries.
- Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority.
- Strong communication, prioritization, and time-management skills across multiple partner engagements.
- Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.
Benefits
Comp & perks- Competitive salaries
- Generous benefits package
- Equity