Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Anyone AI

Research Scientist – Remote, US, LATAM

Anyone AI

Research Scientist designing AI evaluation benchmarks for ML capabilities at Anyone AI Labs. Collaborate with experts to establish rigorous evaluation methods and delivery for lab requests.

Posted 7/6/2026contractRemote • 💃 Anywhere in Latin AmericaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in ML evaluation and benchmarking, with a focus on deep LLM benchmarking and the ability to develop rigorous evaluation designs. Proven track record in managing expert pools and ensuring high standards in evaluation processes.

Highest-signal resume keywords
ML Evaluation ResearchDeep LLM Benchmarking ExpertiseEvaluation Design DevelopmentTeam Management and CalibrationFluent English

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Evaluation ResearchBenchmark DevelopmentConstruct ValidityDiscriminationHeadroom MeasurementContamination AnalysisRubric DevelopmentDeterministic VerifiersExpert CalibrationSample Package Delivery
Soft Skills
Team LeadershipCommunication
Industry Keywords
Public BenchmarksEvaluation TargetsSubject-Matter ExpertsRigorous QCLab RelationshipsTechnical Point of ContactPilot ManagementExpert Pool Management

About the role

Key responsibilities & impact
  • Evaluation research. Turn public benchmarks and eval targets into original evaluation designs. Own the hard questions: construct validity, discrimination, headroom, and contamination.
  • Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).
  • Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.
  • Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.
  • Delivery. Turn lab requests into winning sample packages, then own pilots end to end. Nothing ships before it's lab-ready.

Requirements

What you’ll need
  • Research background in ML evaluation or benchmarking — published/open benchmarks, eval research, or equivalent hands-on work labs relied on.
  • Deep LLM benchmarking expertise, with real strength in code-model evaluation.
  • Fluency with how frontier models are measured: rubrics, pass rates, headroom, contamination, and what makes a task discriminate a model.
  • Proven ability to hold a team or expert pool to a rigorous standard.
  • Fluent English. Spanish a nice to have.

Benefits

Comp & perks
  • Health insurance
  • Professional development opportunities