FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer, Forge
Mistral AIResearch Engineer building Forge workflows for Mistral’s full-stack AI solutions. Improving LLM post-training, evaluation, data pipelines, and scalable deployment systems.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates strong Python engineering skills and hands-on experience with PyTorch or JAX, focusing on building and improving ML workflows, including training, evaluation, and deployment. Capable of collaborating with cross-functional teams to enhance system performance and reliability in fast-paced environments.
Highest-signal resume keywords
Python EngineeringML Workflow DevelopmentDistributed TrainingPost-Training EvaluationClear Communication
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonPyTorchJAXLLM TrainingDebuggingData PipelinesTestingCode ReviewOperational OwnershipInfrastructure Fundamentals
Soft Skills
Clear CommunicationHigh AgencyLow Ego
Tools & Technologies
FSDPDeepSpeedMegatronSLURMRayKubernetesKueueKarpenterSkypilot
Industry Keywords
Model AdaptationSynthetic Data GenerationLarge-Scale ML SystemsObservabilityReproducibilityMaintainable AbstractionsScalable SystemsBottleneck Identification
Tech Stack
Tools & technologiesCloudKubernetesPythonPyTorchRay
About the role
Key responsibilities & impact- Turn real customer requirements into reliable training and deployment workflows
- Work end-to-end across model adaptation and post-training, evaluation, data, and infrastructure
- Build and improve post-training and evaluation workflows for CPT, SFT, RL, and distillation
- Develop tools and pipelines for synthetic data generation, data curation, training, evaluation, and deployment
- Debug and harden large-scale ML systems, including distributed training, scheduling/execution, checkpointing, observability, and reproducibility
- Improve the Forge codebase through clear APIs, tests, documentation, and maintainable abstractions
- Advance the RL training stack, including high-throughput asynchronous rollout and scalable post-training systems
- Ensure Forge deployment is seamless and adaptable across diverse clients, hardware, software stacks, cloud, and on-premises environments
- Partner with researchers and infrastructure engineers to translate bottlenecks into concrete system improvements
- Collaborate with scientists, engineers, product, and customer-facing teams to ship maintainable and trusted Forge projects
Requirements
What you’ll need- Strong Python engineering skills and experience working in large codebases, including testing, code review, CI, and operational ownership
- Hands-on experience with PyTorch, JAX, or similar
- Strong systems and infrastructure fundamentals
- Experience with LLM training or post-training, including fine-tuning, RL, distillation, evaluation, and/or data pipelines
- Excellent debugging skills in distributed jobs, data issues, quality regressions, and infrastructure failures
- Clear communication with technical and non-technical stakeholders
- High agency, low ego, and comfort in fast-moving, under-specified environments
- Nice to have: distributed training experience with FSDP, DeepSpeed, Megatron, or similar
- Nice to have: cluster/orchestration experience with SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, or similar
- Nice to have: experience building reliable ML infrastructure, evaluation systems, or large-scale data processing pipelines
- Nice to have: research experience in LLMs, agents, multimodal models, reasoning, code, or domain adaptation
- Nice to have: open-source contributions, publications, or widely used internal tooling
- Nice to have: experience training multi-billion-parameter models and on petabyte- and exabyte-scale datasets
- Nice to have: ability to identify bottlenecks across the stack and drive improvements from first principles
Benefits
Comp & perks- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal allowances
- Transportation allowances
- Other location-specific perks