FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer, Applied AI
HeyMilo AIResearch Engineer building reproducible RL environments and evaluation systems for HeyMilo’s conversational AI hiring workflows. Testing and improving proprietary and open-weight models across HR and recruiting.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing reinforcement learning environments and evaluating model performance within HR and recruiting workflows. Proficient in Python programming, LLM fine-tuning, and creating reproducible experiments in cloud and Docker environments.
Highest-signal resume keywords
PhD In AI Or Machine LearningReinforcement Learning ExpertisePython Software EngineeringLLM Fine-Tuning ExperienceDocker And Cloud Environments
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reinforcement LearningReward DesignPolicy OptimizationEvaluation MethodologyModel TrainingEvaluation FrameworksOpen-Weight ModelsSimulated EnvironmentsTask GenerationRecord Matching
Soft Skills
Problem OwnershipAdaptabilityCollaboration
Tools & Technologies
DockerLinuxCloud Environments
Certifications & Qualifications
PhD In AI Or Machine Learning
Industry Keywords
HR TechnologyApplicant Tracking SystemsRecruiting WorkflowsStaffing Operations
Tech Stack
Tools & technologiesCloudDockerLinuxPython
About the role
Key responsibilities & impact- Design and build reinforcement learning environments for HR and recruiting workflows, including realistic simulators, tool interfaces, and reproducible episodic task generation
- Evaluate model performance across the HR/recruiting stack, including multi-step recruiter workflows, integrations, and record matching
- Design verifiable reward functions for intermediate actions and end states that resist reward hacking
- Build and maintain evaluation harnesses for frontier and open-weight models, including scoring and cost tracking
- Run post-training experiments such as RLVR-style fine-tuning of open models
- Package environments and results for reproducibility and contribute to research write-ups and published evaluations
- Collaborate with analysts to turn task and rubric ground truth into running environments
- Take workflows from problem definition through reproducible environments for model training and evaluation
Requirements
What you’ll need- PhD in AI, Machine Learning, or a closely related field (required)
- Solid grounding in reinforcement learning and LLM post-training, including reward design, policy optimization, and evaluation methodology
- Strong software engineering skills in Python
- Hands-on experience with LLMs, including evaluations, agentic loops, tool calling, and fine-tuning
- Comfortable with Docker, Linux, and cloud environments for reproducible experiments
- Ability to operate in ambiguity, move quickly, and own problems end to end
- Published research or open-source contributions in ML, RL, or evaluation (bonus)
- Experience with RL/evaluation frameworks and simulated or sandboxed environments (bonus)
- Experience training or fine-tuning open-weight models at any scale (bonus)
- Familiarity with HR technology, applicant tracking systems, recruiting workflows, or staffing operations (bonus)
Benefits
Comp & perks- Competitive pay and benefits
- Work reviewed by and published alongside experienced researchers
- High visibility, fast-paced, execution-driven environment
- Ground-floor role with real influence over how the applied AI team builds