Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Giotto.ai

Senior Research Engineer / Research Scientist – Post-Training, Reinforcement Learning & Training Systems

Giotto.ai

Senior Research Engineer/Scientist at Giotto.ai responsible for post-training AI systems. Designing and scaling training and optimisation solutions in a hybrid role based in Switzerland.

Posted 7/30/2026full-timeLausanne • 🇨🇭 SwitzerlandSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in large-scale language model training, including proficiency in Python and PyTorch, with a focus on reinforcement learning and training stability. Capable of designing and executing complex training strategies while ensuring reproducibility and reliability in research software.

Highest-signal resume keywords
Large-Scale Language-Model TrainingPython ProgrammingPyTorch ProficiencyReinforcement LearningExperimental Design

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SFTPreference OptimisationReward ModellingDistributed ExecutionMixed PrecisionOptimisationDebuggingNumerical OptimisationTraining StabilityCheckpointing
Soft Skills
OwnershipCollaborationProblem-SolvingAttention to DetailCommunication
Tools & Technologies
PyTorch DistributedFSDPDeepSpeedMegatron-CoreHigh-Throughput Inference Systems
Industry Keywords
Reinforcement Learning with Verifiable RewardsMulti-Stage CurriculaScalable Rollout-Generation SystemsTraining StrategiesNumerical Instability

Tech Stack

Tools & technologies
Node.jsPythonPyTorch

About the role

Key responsibilities & impact
  • Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate
  • Design and execute full-parameter and parameter-efficient SFT
  • Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods
  • Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long-horizon agent tasks
  • Integrate reward models, verifiers, critics, graders, and process- or outcome-based rewards
  • Build scalable rollout-generation systems for iterative and on-policy training
  • Design multi-stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation
  • Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism
  • Select sharding, precision, checkpointing, optimiser, batch-size, sequence-length, and activation-recomputation strategies
  • Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs
  • Profile and improve accelerator utilisation, communication efficiency, data loading, and end-to-end training time
  • Diagnose numerical instability, communication failures, out-of-memory errors, stragglers, checkpoint issues, and convergence regressions
  • Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting
  • Build reliable checkpointing, recovery, monitoring, and reproducibility procedures
  • Collaborate closely with data, evaluation, infrastructure, and inference teams
  • Contribute clean, tested code, technical reports, and operational runbooks

Requirements

What you’ll need
  • Ownership of large-scale language-model training or post-training runs across multiple machines and accelerators
  • Experience with workloads for which straightforward single-node training or pure data parallelism was insufficient
  • Deep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution
  • Practical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron-Core, or an equivalent framework
  • Ability to select parallelism and sharding strategies based on model, sequence, memory, and network constraints
  • Strong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability
  • Experience operating high-throughput inference or rollout systems as part of a training loop
  • Ability to debug across model code, distributed communication, numerical optimisation, data, and infrastructure
  • Strong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts
  • Experience building reliable, observable, and reproducible research software
  • Personal ownership of consequential decisions affecting a substantial training programme

Benefits

Comp & perks
  • Remote work is supported
  • The team gathers approximately one week per month in a Swiss office
  • Exceptional candidates elsewhere in Europe may be considered