Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
HeyMilo AI

Analyst, Applied AI

HeyMilo AI

Applied AI Analyst designing tasks, datasets, and rubrics to evaluate conversational AI systems. Supporting HeyMilo’s AI-powered hiring workflows through defensible model analysis and published evaluations.

Posted 8/5/2026full-timeColombo • 🇱🇰 Sri LankaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and evaluating AI systems through realistic task creation and data analysis, with a strong foundation in AI, Machine Learning, and statistical methodologies. Proficient in Python for dataset building and analysis, with exceptional attention to detail and clear communication skills.

Highest-signal resume keywords
Master's Or PhD In AIStrong Understanding Of Large Language ModelsHands-On Experience Assessing LLM OutputsProficient In Python And Data ToolingExperience Contributing To Published Research

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AIMachine LearningStatisticsData AnalysisDataset BuildingEvaluation TechniquesBenchmarkingReinforcement Learning ConceptsGrading Rubric AuthoringFailure Mode Analysis
Soft Skills
Attention To DetailIntellectual HonestyClear Written Communication
Industry Keywords
RecruitingStaffingOperations-Heavy Industries

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Design realistic, multi-step tasks that test AI systems on real-world workflows and author grading rubrics that score them
  • Build and curate ground-truth datasets from real operational data and domain research
  • Review model attempts against tasks, analyze failure modes, catch scoring loopholes, and identify where models pass for the wrong reasons
  • Calibrate task difficulty so results separate models meaningfully
  • Analyze score distributions across models and turn results into clear, defensible findings
  • Contribute to research write-ups and published evaluations alongside engineers and academic reviewers

Requirements

What you’ll need
  • Master's or PhD in AI, Machine Learning, Computer Science, Statistics, or a closely related quantitative field
  • Strong understanding of how large language models are trained and evaluated
  • Hands-on experience assessing LLM outputs through evals, annotation, red-teaming, or benchmark work
  • Comfortable with Python and data tooling for building datasets and analyzing results
  • Exceptional attention to detail and intellectual honesty
  • Clear written communication, including explaining subtle failure modes to technical and non-technical readers
  • Familiarity with reinforcement learning concepts such as reward design and verifiable rewards is a bonus
  • Experience contributing to published research or open evaluation suites is a bonus
  • Domain knowledge in recruiting, staffing, or other operations-heavy industries is a bonus

Benefits

Comp & perks
  • Competitive pay and benefits
  • Ground-floor role on a new applied AI team with direct ownership of what we measure and publish
  • Work reviewed by and published alongside experienced researchers
  • High visibility, fast-paced, execution-driven environment