FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Member of Technical Staff, ML Engineer
SuperintelligenceMember of Technical Staff focused on ML training and inference systems for an AI startup. Building infrastructure for research to achieve breakthroughs in physics.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating machine learning training and inference infrastructure, with a strong focus on distributed training, model-serving systems, and cloud infrastructure. Proficient in developing internal platform tools and collaborating effectively with AI researchers to enhance production systems.
Highest-signal resume keywords
Distributed Training (Multi-GPU, Multi-Node)Model-Serving Systems (vLLM, SGLang, Triton)Cloud Infrastructure (GCP, AWS)Infrastructure as Code (Terraform)GPU Infrastructure, CUDA
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning InfrastructureSoftware Engineering FundamentalsTraining LoopsReward SignalsInference-Time BehaviorTraining-as-a-Service APIsInference GatewaysJob SchedulersPerformance EngineeringDebugging and Hardening Systems
Soft Skills
CollaborationProblem-SolvingHands-On Development
Industry Keywords
AI ResearchProduction SystemsCapacity PlanningObservabilityReliability
Tech Stack
Tools & technologiesAWSCloudGoogle Cloud PlatformNode.jsPyTorchRayTerraform
About the role
Key responsibilities & impact- Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.
- Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.
- Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
- Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else's ticket.
- Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.
Requirements
What you’ll need- Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
- Hands-on experience with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).
- Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.
- Enough ML fluency to work productively with AI researchers: you understand training loops, reward signals, and inference-time behavior well enough to debug them, even without designing the algorithms yourself.
- Experience building internal platform tools such as training-as-a-service APIs, inference gateways, or job schedulers.
- Background in GPU infrastructure, CUDA, or performance engineering for ML workloads.
- Experience with cloud infrastructure (GCP, AWS) and infrastructure as code (Terraform or comparable).
- Prior work embedded alongside a research team, turning research code into production systems.
Benefits
Comp & perks- competitive compensation including salary, benefits, and meaningful early-stage equity