FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer – Applied AI
HeyMilo AIResearch Engineer building reproducible reinforcement-learning environments, reward functions, and evaluation harnesses. Improving conversational AI interviewers through model evaluation and post-training experiments at HeyMilo.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing reinforcement learning environments and reward functions, with strong software engineering skills in Python and hands-on experience with LLMs. Capable of conducting evaluations and fine-tuning models while contributing to research and publications.
Highest-signal resume keywords
Reinforcement LearningPython ProgrammingLLM Fine-TuningDockerEvaluation Methodology
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reinforcement LearningReward DesignPolicy OptimizationEvaluation MethodologyPython ProgrammingLLM Post-TrainingAgentic LoopsFine-TuningEvaluation FrameworksSimulated Environments
Soft Skills
Problem OwnershipAdaptability
Tools & Technologies
DockerLinuxCloud Environments
Certifications & Qualifications
Master's or PhD in AIMachine LearningComputer Science
Industry Keywords
Machine LearningReinforcement LearningEvaluationOpen-Weight ModelsResearch Publications
Tech Stack
Tools & technologiesCloudDockerLinuxPython
About the role
Key responsibilities & impact- Design and build reinforcement learning environments for specific real-world use cases, including realistic simulators, tool interfaces, and episodic task generation
- Design verifiable reward functions that score intermediate actions and end states while resisting reward hacking
- Build and maintain evaluation harnesses for task suites across frontier and open-weight models, including scoring and cost tracking
- Run post-training experiments such as RLVR-style fine-tuning of open models
- Package environments and results for reproducibility
- Contribute to research write-ups and published evaluations
- Help define future use cases based on model weaknesses
- Work directly with founders and research advisors from problem definition through reproducible environments for model evaluation and training
Requirements
What you’ll need- Master's or PhD in AI, Machine Learning, Computer Science, or a closely related field
- Solid grounding in reinforcement learning and LLM post-training, including reward design, policy optimization, and evaluation methodology
- Strong software engineering skills in Python
- Hands-on experience with LLMs, including running evaluations, building agentic loops, tool calling, and fine-tuning
- Comfortable with containers and infrastructure including Docker, Linux, and cloud environments
- Ability to operate in ambiguity and move quickly while owning a problem end to end
- Based in or willing to relocate to the San Francisco Bay Area
- Published research or open-source contributions in ML, RL, or evaluation preferred as a bonus
- Experience with RL/evaluation frameworks and simulated or sandboxed environments is a bonus
- Experience training or fine-tuning open-weight models at any scale is a bonus
Benefits
Comp & perks- Equity
- Benefits
- Opportunity to work directly with founders and experienced research advisors
- Name on published work
- Ground-floor role with real influence over how the company builds