Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Together AI

Technical Support Engineer – Inference

Together AI

Technical Support Engineer maintaining Together AI’s GPU-based inference infrastructure and customer endpoints. Resolving complex AI service incidents while supporting Kubernetes, observability, and infrastructure-as-code operations on US weekend shifts.

Posted 8/4/2026full-timeRemote • 🇺🇸 United StatesMid-LevelSenior💰 $160,000 - $230,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing GPU clusters and AI services, with strong capabilities in Kubernetes, Infrastructure as Code, and complex problem-solving. Proficient in customer communication and cross-functional collaboration to ensure system health and performance.

Highest-signal resume keywords
Kubernetes ManagementAI and ML TechnologiesInfrastructure as CodePython and TypeScript ProgrammingPrometheus and Grafana Expertise

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesPythonTypeScriptAnsibleSLURMGPU TechnologiesHPC EnvironmentsREST API DebuggingInfrastructure as CodeLLM Inference Frameworks
Soft Skills
Excellent CommunicationComplex Problem-SolvingInterpersonal SkillsProject ManagementCross-Functional Collaboration
Tools & Technologies
AWSGCPAzureVast StorageWeka StorageNFS StorageHigh-Performance Network FabricsCurlPostmanGrafana
Industry Keywords
SREDevOpsAI ServicesGPU Cluster ManagementCompute-Cluster AdministrationTechnical SupportIncident ManagementSystem Health MonitoringTraffic RoutingDocumentation

Tech Stack

Tools & technologies
AnsibleAWSAzureGoogle Cloud PlatformGrafanaJavaScriptKubernetesNFSPrometheusPythonSwitchingTypeScript

About the role

Key responsibilities & impact
  • Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference, and fine-tuning services
  • Act as a customer-facing SRE to keep customer inference endpoints on Kubernetes healthy, stable, and performant
  • Become a product expert and serve as the last technical defense before escalation to Engineering and Product
  • Assist with hardware and platform migrations by validating system health and traffic routing
  • Monitor dashboards, detect anomalies, and escalate with data-backed analysis
  • Manage customer communications during incidents and degradations
  • Translate technical findings into clear, evidence-backed customer updates
  • Execute infrastructure changes through pull requests and infrastructure-as-code for endpoint configuration, model deployment, capacity scaling, and cluster configuration
  • Flag engine-level bugs with logs and reproduction steps
  • Collaborate with Engineering, Research, Product, Sales, Support, and senior leaders to address customer concerns
  • Identify support-case patterns and help drive Together AI’s product roadmap
  • Maintain documentation covering system configurations, procedures, troubleshooting guides, and FAQs
  • Provide support coverage during holidays, nights, and weekends as required

Requirements

What you’ll need
  • 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, including at least 1 year supporting an AI service
  • Experience as an SRE or DevOps engineer working with Kubernetes
  • Strong knowledge of AI, ML, GPU technologies, and HPC environments
  • Production-level experience with Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS storage, and container infrastructure
  • Familiarity with Vast and Weka storage systems in HPC environments
  • Ability to diagnose complex network-layer issues and read traces
  • Strong knowledge of Python, TypeScript, and/or JavaScript, with curl and Postman-like testing/debugging experience
  • Expertise with Prometheus and Grafana at scale
  • Deep familiarity with REST API debugging and HTTP semantics
  • Experience with LLM inference frameworks, LoRA fine-tuning, and common training failure modes
  • Experience with Infrastructure as Code and Git-based workflows
  • Background in GPU cluster management
  • Experience with AWS, GCP, and/or Azure
  • Understanding of compute-cluster installation, configuration, administration, troubleshooting, and security
  • Complex technical problem-solving and troubleshooting ability
  • Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research
  • Strong ownership and willingness to learn
  • Excellent communication and interpersonal skills, including explaining complex technical concepts to non-technical stakeholders
  • Ability to manage multiple projects, context switching, and prioritization
  • Ability to work US daytime hours, including Saturday and Sunday plus two additional weekdays
  • Ability to work a four-day, ten-hour shift with additional weekend on-call coverage
  • Flexibility to provide holiday, night, and weekend support

Benefits

Comp & perks
  • Startup equity
  • Health insurance
  • Other benefits
  • Flexibility in terms of remote work
  • Competitive compensation