Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
InnoData

AI/ML Research Engineer, LLM Post-Training – Evaluation

InnoData

AI/ML Research Engineer designing and implementing LLM training and evaluation pipelines. Collaborating with technical teams to improve foundation model performance at Innodata.

Posted 7/27/2026full-timeRemote • 🇺🇸 United StatesJuniorMid-Level💰 $80,000 - $175,000 per yearWebsite

Tech Stack

Tools & technologies
PythonPyTorchTensorflow

About the role

Key responsibilities & impact
  • design and implement the pipelines and tooling that connect data, evaluation, and post-training
  • help customers and internal teams move from evaluation findings to measurable model improvements
  • build fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization)
  • integrate evaluation harnesses into model development loops
  • improve experiment reliability and throughput
  • support advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions
  • contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation.
  • Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery
  • Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking
  • Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses
  • Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows
  • Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring
  • Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift
  • Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems
  • Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions
  • Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements
  • Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects
  • Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team.

Requirements

What you’ll need
  • BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field (MS/PhD preferred)
  • 2-3 years of relevant industry or research engineering experience in ML/AI systems
  • Hands-on experience with LLM training / fine-tuning / post-training, including at least one of:
  • supervised fine-tuning (SFT)
  • preference optimization (e.g., DPO or related methods)
  • RLHF / RLAIF-style workflows
  • task- or domain-adaptation of foundation models
  • Strong programming skills in Python and experience building production-quality ML code
  • Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks)
  • Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons
  • Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging
  • Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads (GPU/accelerator environments preferred)
  • Experience with large-scale data processing and workflow orchestration in support of model training/evaluation
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads
  • Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences.

Benefits

Comp & perks
  • 🌐 Worldwide ❌ Jobs You've Hidden ⭐️ Saved Jobs ✅ Applied Jobs ✉️ Email Alerts 👤 Account InnoData Website LinkedIn All Job Openings 2 - 10 employees Founded 2019 🤝 B2B 💼 Consulting 🌍 Social Impact B2B
  • Consulting
  • Social Impact InnoData is INNOvation DATA SCS, an Italian social cooperative based in Foggia that identifies itself as a provider of technological solutions. The company website (currently under maintenance) highlights "Soluzioni tecnologiche" (technological solutions) and emphasizes social impact ("Impatto sociale"). Contact details listed include Via Francesco Crispi 65, 71121 Foggia, Italy. Based on the available information, InnoData appears to operate at the intersection of technology and social impact, likely offering tech-focused services to other organizations. AI/ML Research Engineer, LLM Post-Training – Evaluation Job not on LinkedIn 🔥 27 minutes ago 🇺🇸 United States – Remote 💵 $80k - $175k / year ⏰ Full Time 🟢 Junior 🟡 Mid-level 🗣️ LLM Engineer Python PyTorch Tensorflow Apply Now Find Hiring Managers Customize resume + cover letter Report problem ☆ Save ☑️ Mark as applied ❌ Hide 📋 Description
  • design and implement the pipelines and tooling that connect data, evaluation, and post-training
  • help customers and internal teams move from evaluation findings to measurable model improvements
  • build fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization)
  • integrate evaluation harnesses into model development loops
  • improve experiment reliability and throughput
  • support advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions
  • contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation.
  • Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery
  • Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking
  • Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses
  • Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows
  • Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring
  • Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift
  • Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems
  • Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions
  • Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements
  • Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects
  • Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team. 🎯 Requirements
  • BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field (MS/PhD preferred)
  • 2-3 years of relevant industry or research engineering experience in ML/AI systems
  • Hands-on experience with LLM training / fine-tuning / post-training, including at least one of:
  • supervised fine-tuning (SFT)
  • preference optimization (e.g., DPO or related methods)
  • RLHF / RLAIF-style workflows
  • task- or domain-adaptation of foundation models
  • Strong programming skills in Python and experience building production-quality ML code
  • Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks)
  • Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons
  • Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging
  • Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads (GPU/accelerator environments preferred)
  • Experience with large-scale data processing and workflow orchestration in support of model training/evaluation
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads
  • Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences. Apply Now 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score Similar Jobs AI Infrastructure Co-Founder / CEO 🔥 3 hours ago EWOR 201 - 500 📚 Education 💸 Finance 💼 Consulting Website LinkedIn All Job Openings Co-Founder / CEO responsible for launching and scaling AI Infrastructure startups in a supportive environment. Engage with experienced entrepreneurs while developing a high-growth venture. 🇺🇸 United States – Remote 💰 $34.2M Series A - EWOR on 2025-04 ⏰ Full Time 🟡 Mid-level 🟠 Senior 🗣️ LLM Engineer AI Infrastructure Co-Founder, CCO 🔥 3 hours ago EWOR 201 - 500 📚 Education 💸 Finance 💼 Consulting Website LinkedIn All Job Openings Co-Founder / CCO leading an AI Infrastructure startup backed by experienced entrepreneurs. Building startup with coaching and funding support to reach significant revenues. 🇺🇸 United States – Remote 💰 $34.2M Series A - EWOR on 2025-04 ⏰ Full Time 🟡 Mid-level 🟠 Senior 🗣️ LLM Engineer Forward Deployed Engineer, Generative AI 🕒 July 14 Tiger Analytics 1001 - 5000 🏥 Healthcare 📦 Logistics 📣 Marketing Website LinkedIn All Job Openings Forward Deployed Engineer at Tiger Analytics driving deployment and scaling of Generative AI solutions. Collaborate with teams to operationalize AI research into production-grade infrastructure. 🇺🇸 United States – Remote ⏰ Full Time 🟡 Mid-level 🟠 Senior 🗣️ LLM Engineer 🦅 H1B Visa Sponsor AWS Azure BigQuery Cloud Google Cloud Platform Python SQL Account Executive, AI Infrastructure Sales 🕒 June 29 Vultr 201 - 500 🤖 Artificial Intelligence 🤝 B2B 🔧 Hardware Website LinkedIn All Job Openings Account Executive at Vultr responsible for driving growth in AI Infrastructure Sales. Own strategic customer relationships and guide clients through AI cloud infrastructure adoption. 🇺🇸 United States – Remote 💵 $90k - $110k / year 💰 $329M Debt Financing - Vultr on 2025-06 ⏰ Full Time 🟡 Mid-level 🟠 Senior 🗣️ LLM Engineer Cloud Research Engineer 5 – LLM-Driven Product Understanding 🕒 June 27 Netflix 10,000+ employees 📱 Media 👥 B2C Website LinkedIn All Job Openings Research Engineer at Netflix developing LLM prototypes to enhance member understanding. Collaborating with cross-functional teams in a dynamic and innovative environment. 🇺🇸 United States – Remote 💵 $466k - $750k / year ⏰ Full Time 🟡 Mid-level 🟠 Senior 🗣️ LLM Engineer Python PyTorch Tensorflow View More LLM Engineer Jobs 🌐 Worldwide Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or support@remoterocketship.com Search Search Jobs by country Search jobs by city Search jobs by job title Search entry-level jobs Search junior-level jobs Search senior-level jobs Search jobs by tech stack Search jobs by contract type Search remote internships Search remote part-time jobs Remote jobs Anywhere in the World Companies Hiring Anywhere in the World Companies Hiring Sales People Anywhere in the World Companies Hiring Software Engineers Anywhere in the World Resources Advice Tips for finding remote jobs Interview questions and answers Resume examples Cover letter examples Post a job Affiliates Is Remote Rocketship legit? Privacy policy Terms of service Job board SEO course OpenClaw job finder Find jobs using your resume Jobs by Country Remote jobs anywhere in the world (Worldwide remote jobs) Remote jobs United States Remote jobs Australia Remote jobs Brazil Remote jobs Canada Remote jobs France Remote jobs Ireland Remote jobs Germany Remote jobs Netherlands Remote jobs Spain Remote jobs UK Popular Jobs Remote data analyst jobs Remote customer support jobs Remote executive assistant jobs Remote marketing jobs Remote product designer jobs Remote product manager jobs Remote project manager jobs Remote recruiter jobs Remote sales jobs Remote software engineer jobs Jobs by Type Remote full-time jobs Remote part-time jobs Remote contract jobs Remote internship jobs Remote entry-level jobs Remote jobs with no experience required Remote junior jobs (1-3 years of experience) Digital nomad jobs Remote jobs with no degree required Freelance remote jobs Temporary remote jobs Remote jobs hiring now Stay at home mom jobs