FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer, ML Platform
Mistral AIResearch Engineer building Mistral AI’s distributed GPU platform for large-scale training, evaluation, and inference. Improving Kubernetes orchestration, reliability, and researcher self-service workflows.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and optimizing ML infrastructure and distributed systems, with a strong focus on Kubernetes, GPU resource management, and performance diagnostics. Proficient in creating self-service workflows and operational tooling to enhance the efficiency of critical ML workloads.
Highest-signal resume keywords
Kubernetes Platform EngineeringPython or Go ProficiencyDistributed Systems ExperienceGPU Infrastructure FamiliarityML Workload Understanding
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonGoKubernetesDistributed SystemsGPU Resource ManagementPyTorchCUDANCCLObservabilityCapacity Planning
Soft Skills
Problem DiagnosisDeveloper Experience FocusAdaptability
Tools & Technologies
KueueKarpenterVolcanoKyverno
Industry Keywords
ML InfrastructureBatch InferenceWorkload AdmissionTopology AwarenessScheduling Latency
Tech Stack
Tools & technologiesDistributed SystemsGoKubernetesPythonPyTorch
About the role
Key responsibilities & impact- Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference
- Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement of GPU workloads
- Improve provisioning, allocation, and utilization of heterogeneous GPU resources across clusters
- Place workloads across multiple clusters based on capacity, data locality, hardware requirements, and organizational priorities
- Create self-service workflows for launching, observing, debugging, and reproducing distributed workloads
- Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency
- Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads
- Participate in on-call rotations and troubleshoot applications, schedulers, networking, storage, and GPU infrastructure
Requirements
What you’ll need- 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field
- Proficiency in Python or Go
- Experience working with production-grade distributed systems
- Strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management
- Understanding of Kueue, Karpenter, Volcano, and Kyverno
- Understanding of distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference
- Familiarity with GPU infrastructure, PyTorch, CUDA, NCCL, and high-performance networking
- Understanding of quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission
- Ability to diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware
- Interest in developer experience and turning complex infrastructure into simple, reliable interfaces
- Ability to thrive in an ambiguous, fast-moving environment shaped by frontier AI research
- Willingness to commute to Palo Alto or San Francisco offices 4 days a week
Benefits
Comp & perks- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal allowances
- Transportation allowances
- Other location-specific perks