FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Research Engineer / Research Scientist – Post-Training, Reinforcement Learning & Training Systems
Giotto.aiSenior Research Engineer/Scientist at Giotto.ai responsible for post-training AI systems. Designing and scaling training and optimisation solutions in a hybrid role based in Switzerland.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in large-scale language model training, including proficiency in Python and PyTorch, with a focus on reinforcement learning and training stability. Capable of designing and executing complex training strategies while ensuring reproducibility and reliability in research software.
Highest-signal resume keywords
Large-Scale Language-Model TrainingPython ProgrammingPyTorch ProficiencyReinforcement LearningExperimental Design
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SFTPreference OptimisationReward ModellingDistributed ExecutionMixed PrecisionOptimisationDebuggingNumerical OptimisationTraining StabilityCheckpointing
Soft Skills
OwnershipCollaborationProblem-SolvingAttention to DetailCommunication
Tools & Technologies
PyTorch DistributedFSDPDeepSpeedMegatron-CoreHigh-Throughput Inference Systems
Industry Keywords
Reinforcement Learning with Verifiable RewardsMulti-Stage CurriculaScalable Rollout-Generation SystemsTraining StrategiesNumerical Instability
Tech Stack
Tools & technologiesNode.jsPythonPyTorch
About the role
Key responsibilities & impact- Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate
- Design and execute full-parameter and parameter-efficient SFT
- Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods
- Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long-horizon agent tasks
- Integrate reward models, verifiers, critics, graders, and process- or outcome-based rewards
- Build scalable rollout-generation systems for iterative and on-policy training
- Design multi-stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation
- Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism
- Select sharding, precision, checkpointing, optimiser, batch-size, sequence-length, and activation-recomputation strategies
- Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs
- Profile and improve accelerator utilisation, communication efficiency, data loading, and end-to-end training time
- Diagnose numerical instability, communication failures, out-of-memory errors, stragglers, checkpoint issues, and convergence regressions
- Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting
- Build reliable checkpointing, recovery, monitoring, and reproducibility procedures
- Collaborate closely with data, evaluation, infrastructure, and inference teams
- Contribute clean, tested code, technical reports, and operational runbooks
Requirements
What you’ll need- Ownership of large-scale language-model training or post-training runs across multiple machines and accelerators
- Experience with workloads for which straightforward single-node training or pure data parallelism was insufficient
- Deep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution
- Practical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron-Core, or an equivalent framework
- Ability to select parallelism and sharding strategies based on model, sequence, memory, and network constraints
- Strong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability
- Experience operating high-throughput inference or rollout systems as part of a training loop
- Ability to debug across model code, distributed communication, numerical optimisation, data, and infrastructure
- Strong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts
- Experience building reliable, observable, and reproducible research software
- Personal ownership of consequential decisions affecting a substantial training programme
Benefits
Comp & perks- Remote work is supported
- The team gathers approximately one week per month in a Swiss office
- Exceptional candidates elsewhere in Europe may be considered