Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Johnson & Johnson

Postdoctoral Scholar, AI Evaluation & Standards

Johnson & Johnson

Postdoctoral researcher at Johnson & Johnson designing evaluation methods for GenAI tools in pharmaceutical R&D contexts. Roles involve building benchmarks, running evaluations, and translating findings into improvements.

Posted 8/1/2026full-timeBarcelona • New Jersey • 🇺🇸 United StatesMid-LevelSenior💰 €43,600 - €70,150 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing evaluation frameworks and criteria for GenAI tools in pharmaceutical R&D, with a strong foundation in biomedical science and data analysis. Proficient in translating expert judgment into measurable outcomes and improving evaluation methods through collaboration with cross-functional teams.

Highest-signal resume keywords
PhD In Biomedical ScienceExperience In Evaluation MethodsProficiency In PythonUnderstanding Of Pharmaceutical R&DAbility To Analyze Model Outputs

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Evaluation Framework DesignBenchmark Dataset DevelopmentAnnotation ProtocolsQuality ReviewsData AnalysisFailure Pattern AnalysisGenAI TestingRubric DevelopmentClinical ResearchTranslational Science
Soft Skills
Clear Written CommunicationVerbal Communication
Tools & Technologies
Data Science ToolsMachine Learning Tools
Industry Keywords
Biomedical SciencePharmaceutical R&DRegulatory ScienceClinical DevelopmentBiomedical InformaticsAI/MLComputational BiologyBioinformaticsTherapeutic Area ScienceTranslational Science

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
  • Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
  • Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
  • Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
  • Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
  • Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
  • Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
  • Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
  • Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
  • Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
  • Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.

Requirements

What you’ll need
  • PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
  • Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
  • Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
  • Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
  • Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
  • Proficiency in Python and common data science or machine learning tools.
  • Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
  • Clear written and verbal communication skills.

Benefits

Comp & perks
  • an annual bonus with set target (% of pay) depending on pay grade / location, where the actual amount is based on the employees’ and companies’ performance of the previous calendar year, or sales commissions.
  • vacation days
  • parental leave for a minimum of 12 weeks
  • bereavement leave
  • caregiver leave
  • volunteer leave
  • well-being reimbursement
  • programs for financial, physical and mental health.
  • service anniversary and recognition awards
  • employees - and in some locations - eligible dependents - can participate in several insurance plans