FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Technical Support Engineer – GPU Clusters
Together AITechnical Support Engineer maintaining Kubernetes GPU clusters for Together AI, a research-driven open AI infrastructure company. Resolving enterprise customer infrastructure, networking, and storage issues.
Posted 8/4/2026full-timeRemote • California • 🇺🇸 United StatesMid-LevelSenior💰 $160,000 - $230,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing Kubernetes GPU clusters and high-performance computing environments, with a strong focus on customer engagement and technical problem-solving. Proficient in operating production infrastructure and collaborating cross-functionally to enhance customer satisfaction and drive product improvements.
Highest-signal resume keywords
Kubernetes AdministrationSRE ExperienceAI/ML TechnologiesHPC/Slurm Cluster ManagementComplex Problem-Solving
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesSLURMAnsibleGPU TechnologiesHigh-Performance ComputingScripting LanguagesProgramming LanguagesDistributed Storage SystemsNetwork DiagnosticsContainer Infrastructure
Soft Skills
Excellent CommunicationInterpersonal SkillsOwnershipProactive Issue ResolutionProject Management
Tools & Technologies
NFS StorageInfiniBandRDMAHigh-Performance Network FabricsCompute Clusters
Industry Keywords
Customer-Facing Technical RoleAI Service SupportMission-Critical SaaS APITechnical DocumentationTroubleshooting Guides
Tech Stack
Tools & technologiesAnsibleKubernetesNFSNode.jsSwitching
About the role
Key responsibilities & impact- Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters
- Act as a customer-facing SRE to keep customer Kubernetes clusters healthy and stable
- Become a product expert and serve as the last technical defense before escalation to Engineering and Product
- Monitor GPU cluster health and communicate hardware issues with remediation steps
- Operate and maintain production infrastructure, including fleet rebalancing, Slurm maintenance, node repair/migration, and Kubernetes workload management
- Investigate and resolve storage and networking issues in bare-metal and VM environments
- Collaborate with Engineering, Research, Product, Sales, Support, and senior leaders to address customer concerns and drive customer satisfaction
- Identify support-case patterns and turn customer insights into roadmap improvements
- Maintain documentation, troubleshooting guides, procedures, and FAQs
- Provide holiday, night, and weekend coverage as required
Requirements
What you’ll need- 3+ years of experience in a customer-facing technical role, including at least 1 year supporting an AI service or mission-critical SaaS API
- Experience as an SRE or DevOps engineer working with Kubernetes
- Knowledge of AI, ML, GPU technologies, and HPC environments
- Advanced knowledge of Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS storage, container infrastructure, scripting, and programming languages
- Experience with HPC/Slurm cluster environments, including node draining, job scheduling, and maintenance workflows
- Familiarity with InfiniBand, RDMA, and network interface diagnostics
- Experience with distributed storage systems such as Weka and NFS, including I/O and bandwidth troubleshooting
- Understanding of installing, configuring, administering, troubleshooting, and securing compute clusters
- Complex technical problem-solving and proactive issue resolution
- Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research
- Strong ownership and willingness to learn new skills
- Excellent communication and interpersonal skills, including explaining technical concepts to non-technical stakeholders
- Ability to manage multiple projects, context switching, and prioritization
- Willingness to work US daytime hours, weekends, holidays, nights, and a four-day, 10-hour shift with weekend on-call coverage
Benefits
Comp & perks- Startup equity
- Health insurance
- Other benefits
- Flexibility in terms of remote work