FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Analyst, Applied AI
HeyMilo AIApplied AI Analyst designing tasks, datasets, and rubrics to evaluate conversational AI systems. Supporting HeyMilo’s AI-powered hiring workflows through defensible model analysis and published evaluations.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and evaluating AI systems through realistic task creation and data analysis, with a strong foundation in AI, Machine Learning, and statistical methodologies. Proficient in Python for dataset building and analysis, with exceptional attention to detail and clear communication skills.
Highest-signal resume keywords
Master's Or PhD In AIStrong Understanding Of Large Language ModelsHands-On Experience Assessing LLM OutputsProficient In Python And Data ToolingExperience Contributing To Published Research
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AIMachine LearningStatisticsData AnalysisDataset BuildingEvaluation TechniquesBenchmarkingReinforcement Learning ConceptsGrading Rubric AuthoringFailure Mode Analysis
Soft Skills
Attention To DetailIntellectual HonestyClear Written Communication
Industry Keywords
RecruitingStaffingOperations-Heavy Industries
Tech Stack
Tools & technologiesPython
About the role
Key responsibilities & impact- Design realistic, multi-step tasks that test AI systems on real-world workflows and author grading rubrics that score them
- Build and curate ground-truth datasets from real operational data and domain research
- Review model attempts against tasks, analyze failure modes, catch scoring loopholes, and identify where models pass for the wrong reasons
- Calibrate task difficulty so results separate models meaningfully
- Analyze score distributions across models and turn results into clear, defensible findings
- Contribute to research write-ups and published evaluations alongside engineers and academic reviewers
Requirements
What you’ll need- Master's or PhD in AI, Machine Learning, Computer Science, Statistics, or a closely related quantitative field
- Strong understanding of how large language models are trained and evaluated
- Hands-on experience assessing LLM outputs through evals, annotation, red-teaming, or benchmark work
- Comfortable with Python and data tooling for building datasets and analyzing results
- Exceptional attention to detail and intellectual honesty
- Clear written communication, including explaining subtle failure modes to technical and non-technical readers
- Familiarity with reinforcement learning concepts such as reward design and verifiable rewards is a bonus
- Experience contributing to published research or open evaluation suites is a bonus
- Domain knowledge in recruiting, staffing, or other operations-heavy industries is a bonus
Benefits
Comp & perks- Competitive pay and benefits
- Ground-floor role on a new applied AI team with direct ownership of what we measure and publish
- Work reviewed by and published alongside experienced researchers
- High visibility, fast-paced, execution-driven environment