Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Together AI

Technical Support Engineer – GPU Clusters

Together AI

Technical Support Engineer maintaining Kubernetes GPU clusters for Together AI, a research-driven open AI infrastructure company. Resolving enterprise customer infrastructure, networking, and storage issues.

Posted 8/4/2026full-timeRemote • California • 🇺🇸 United StatesMid-LevelSenior💰 $160,000 - $230,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing Kubernetes GPU clusters and high-performance computing environments, with a strong focus on customer engagement and technical problem-solving. Proficient in operating production infrastructure and collaborating cross-functionally to enhance customer satisfaction and drive product improvements.

Highest-signal resume keywords
Kubernetes AdministrationSRE ExperienceAI/ML TechnologiesHPC/Slurm Cluster ManagementComplex Problem-Solving

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesSLURMAnsibleGPU TechnologiesHigh-Performance ComputingScripting LanguagesProgramming LanguagesDistributed Storage SystemsNetwork DiagnosticsContainer Infrastructure
Soft Skills
Excellent CommunicationInterpersonal SkillsOwnershipProactive Issue ResolutionProject Management
Tools & Technologies
NFS StorageInfiniBandRDMAHigh-Performance Network FabricsCompute Clusters
Industry Keywords
Customer-Facing Technical RoleAI Service SupportMission-Critical SaaS APITechnical DocumentationTroubleshooting Guides

Tech Stack

Tools & technologies
AnsibleKubernetesNFSNode.jsSwitching

About the role

Key responsibilities & impact
  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters
  • Act as a customer-facing SRE to keep customer Kubernetes clusters healthy and stable
  • Become a product expert and serve as the last technical defense before escalation to Engineering and Product
  • Monitor GPU cluster health and communicate hardware issues with remediation steps
  • Operate and maintain production infrastructure, including fleet rebalancing, Slurm maintenance, node repair/migration, and Kubernetes workload management
  • Investigate and resolve storage and networking issues in bare-metal and VM environments
  • Collaborate with Engineering, Research, Product, Sales, Support, and senior leaders to address customer concerns and drive customer satisfaction
  • Identify support-case patterns and turn customer insights into roadmap improvements
  • Maintain documentation, troubleshooting guides, procedures, and FAQs
  • Provide holiday, night, and weekend coverage as required

Requirements

What you’ll need
  • 3+ years of experience in a customer-facing technical role, including at least 1 year supporting an AI service or mission-critical SaaS API
  • Experience as an SRE or DevOps engineer working with Kubernetes
  • Knowledge of AI, ML, GPU technologies, and HPC environments
  • Advanced knowledge of Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS storage, container infrastructure, scripting, and programming languages
  • Experience with HPC/Slurm cluster environments, including node draining, job scheduling, and maintenance workflows
  • Familiarity with InfiniBand, RDMA, and network interface diagnostics
  • Experience with distributed storage systems such as Weka and NFS, including I/O and bandwidth troubleshooting
  • Understanding of installing, configuring, administering, troubleshooting, and securing compute clusters
  • Complex technical problem-solving and proactive issue resolution
  • Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research
  • Strong ownership and willingness to learn new skills
  • Excellent communication and interpersonal skills, including explaining technical concepts to non-technical stakeholders
  • Ability to manage multiple projects, context switching, and prioritization
  • Willingness to work US daytime hours, weekends, holidays, nights, and a four-day, 10-hour shift with weekend on-call coverage

Benefits

Comp & perks
  • Startup equity
  • Health insurance
  • Other benefits
  • Flexibility in terms of remote work