FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Scientist
ScienceLogicData Scientist evaluating local LLMs and agentic systems at ScienceLogic, an AIOps platform company. Building predictive models for operational telemetry, anomalies, and early-warning signals.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and evaluating LLM and agentic systems, with a strong focus on metrics, model calibration, and predictive analytics. Proficient in Python and SQL, with hands-on experience in deploying and monitoring models in production environments.
Highest-signal resume keywords
LLM EvaluationPredictive ModelingApplied StatisticsPython ProficiencySQL Querying
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data ScienceMachine LearningQuantitative AnalysisExperiment DesignModel CalibrationEvaluation MetricsStatistical AnalysisTime-Series ModelsNLP SystemsData Visualization
Soft Skills
CommunicationCollaborationProblem-Solving
Tools & Technologies
Eval/Harness FrameworksJudge PipelinesModel-Serving LibrariesAIOpsObservability Tools
Industry Keywords
Foundation ModelsRed-TeamingIT OperationsBig DataEnterprise Security
Tech Stack
Tools & technologiesCloudITSMPythonSQL
About the role
Key responsibilities & impact- Design and own evaluation harnesses for LLM and agentic outputs, including golden sets, regression suites, and rubric-based scoring
- Build and calibrate LLM-as-judge pipelines and validate judges against human labels
- Define and track response-quality metrics including groundedness, hallucination rate, relevance, completeness, instruction-following, and persona adherence
- Curate, version, and expand evaluation datasets
- Benchmark local models and quantify quality tradeoffs versus larger alternatives
- Red-team systems for prompt injection, jailbreaks, tool misuse, and edge cases
- Design chaos and stress tests for model and agent reliability
- Characterize failure modes and improve guardrails and regression coverage
- Evaluate retrieval quality and experiment with chunking, indexing, and hybrid retrieval
- Analyze multi-step agent trajectories, tool-call correctness, efficiency, replayable state, and guardrail breaches
- Assess intent classification and routing quality
- Monitor behavioral regressions and production output-quality drift
- Recommend fixes through prompts, retrieval, grounding, routing, and model selection
- Build, deploy, monitor, recalibrate, and retrain production models for capacity forecasting, anomaly prediction, and early-warning signals
- Define operational accuracy and lead-time metrics
- Integrate predictive signals into LLM and agentic layers for reasoning, advisories, and recommendations
- Apply AIOps/NOC analysis to log anomalies, event correlation, root-cause analysis, and problem analysis
- Quantify system economics, token consumption, interaction types, MTTR, and operator-hours
- Communicate analysis to engineering and product stakeholders
- Use LLM-assisted workflows, synthetic evaluation cases, and bootstrapped labeled data
- Track and adopt evaluation, retrieval, and agentic-analysis techniques
Requirements
What you’ll need- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field — or equivalent experience
- 3+ years in data science, ML, or applied quantitative analysis
- Strong applied statistics and ability to design sound experiments and significance tests on noisy, non-deterministic outputs
- Experience building, deploying, and monitoring predictive or time-series models in production, including recalibration as data shifts
- Demonstrated work evaluating, analyzing, or improving LLM or NLP systems
- Proficiency in Python
- Strong SQL and comfort querying large analytical datasets
- Fluency with foundation models and hands-on experience with modern LLM evaluation and tooling, including eval/harness frameworks, judge pipelines, and model-serving, prompting, and testing libraries
- Ability to build analysis and visualization in code
- Preferred: experience with small or self-hosted/local models, retrieval-augmented systems, agentic frameworks, red-teaming, IT operations/AIOps/NOC/ITSM/observability, large-scale analytical and big-data stores, cloud data science and ML workloads, and enterprise security and compliance constraints
Benefits
Comp & perks- 🌐 Worldwide ❌ Jobs You've Hidden ⭐️ Saved Jobs ✅ Applied Jobs ✉️ Email Alerts 👤 Account ScienceLogic Website LinkedIn All Job Openings 501 - 1000 employees Founded 2003 🤖 Artificial Intelligence ☁️ SaaS 🤝 B2B 💰 $21.2M Venture Round - ScienceLogic on 2022-10 Artificial Intelligence
- SaaS
- B2B ScienceLogic is a provider of AI-driven observability and AIOps software that helps organizations observe, automate, and troubleshoot complex hybrid IT environments. Its ScienceLogic AI Platform (branded Skylar) combines hybrid cloud monitoring, network and application observability, automated root-cause analysis, workflow automation, compliance and configuration management, and extensive integrations to reduce mean time to repair (MTTR) and consolidate IT tools. The company delivers its platform as a SaaS solution (with rapid deployment options), targets enterprise and service-provider IT operations, and emphasizes AI/automation to drive operational efficiency, resiliency, and cost reduction. Data Scientist 🔥 1 minute ago 🏢🏡 null – Hybrid ⏰ Full Time 🟡 Mid-level 🟠 Senior 📊 Data Scientist 👻 Ghost score 12% Apply Now Customize resume + cover letter Report problem ☆ Save ☑️ Mark as applied ❌ Hide 📋 Description
- Design and own evaluation harnesses for LLM and agentic outputs, including golden sets, regression suites, and rubric-based scoring
- Build and calibrate LLM-as-judge pipelines and validate judges against human labels
- Define and track response-quality metrics including groundedness, hallucination rate, relevance, completeness, instruction-following, and persona adherence
- Curate, version, and expand evaluation datasets
- Benchmark local models and quantify quality tradeoffs versus larger alternatives
- Red-team systems for prompt injection, jailbreaks, tool misuse, and edge cases
- Design chaos and stress tests for model and agent reliability
- Characterize failure modes and improve guardrails and regression coverage
- Evaluate retrieval quality and experiment with chunking, indexing, and hybrid retrieval
- Analyze multi-step agent trajectories, tool-call correctness, efficiency, replayable state, and guardrail breaches
- Assess intent classification and routing quality
- Monitor behavioral regressions and production output-quality drift
- Recommend fixes through prompts, retrieval, grounding, routing, and model selection
- Build, deploy, monitor, recalibrate, and retrain production models for capacity forecasting, anomaly prediction, and early-warning signals
- Define operational accuracy and lead-time metrics
- Integrate predictive signals into LLM and agentic layers for reasoning, advisories, and recommendations
- Apply AIOps/NOC analysis to log anomalies, event correlation, root-cause analysis, and problem analysis
- Quantify system economics, token consumption, interaction types, MTTR, and operator-hours
- Communicate analysis to engineering and product stakeholders
- Use LLM-assisted workflows, synthetic evaluation cases, and bootstrapped labeled data
- Track and adopt evaluation, retrieval, and agentic-analysis techniques 🎯 Requirements
- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field — or equivalent experience
- 3+ years in data science, ML, or applied quantitative analysis
- Strong applied statistics and ability to design sound experiments and significance tests on noisy, non-deterministic outputs
- Experience building, deploying, and monitoring predictive or time-series models in production, including recalibration as data shifts
- Demonstrated work evaluating, analyzing, or improving LLM or NLP systems
- Proficiency in Python
- Strong SQL and comfort querying large analytical datasets
- Fluency with foundation models and hands-on experience with modern LLM evaluation and tooling, including eval/harness frameworks, judge pipelines, and model-serving, prompting, and testing libraries
- Ability to build analysis and visualization in code
- Preferred: experience with small or self-hosted/local models, retrieval-augmented systems, agentic frameworks, red-teaming, IT operations/AIOps/NOC/ITSM/observability, large-scale analytical and big-data stores, cloud data science and ML workloads, and enterprise security and compliance constraints Apply Now 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score Similar Jobs Data Scientist 🔥 5 hours ago EXL 10,000+ employees 🏥 Healthcare 🛡️ Insurance 📦 Logistics Website LinkedIn All Job Openings Data Scientist developing Python-based Generative AI and NLP solutions for client projects. Building advanced RAG pipelines, optimizing models, and integrating AI into scalable systems. 💰 $2M Venture Round on 2015-01 ⏰ Full Time 🟡 Mid-level 🟠 Senior 📊 Data Scientist Engineer III, Data Science 🔥 18 hours ago Verizon 10,000+ employees 📡 Telecommunications 👥 B2C 🏢 Enterprise Website LinkedIn All Job Openings Data scientist developing predictive models, experiments, and analytics for Verizon’s wireless customer growth. Enabling churn reduction, revenue growth, targeting, and campaign optimization. 🏢🏡 Chennai – Hybrid 🔥 Funding within the last year 💰 $2.3G Post IPO debt on 2025-08 ⏰ Full Time 🟡 Mid-level 🟠 Senior 📊 Data Scientist Data Scientist 🕒 Yesterday EXL 10,000+ employees 🏥 Healthcare 🛡️ Insurance 📦 Logistics Website LinkedIn All Job Openings Data Scientist developing scalable Generative AI, NLP, and RAG applications for client solutions. Building models with Python, cloud GPU platforms, and modern LLM technologies in hybrid Noida role. 🏢🏡 Noida – Hybrid 💰 $2M Venture Round on 2015-01 ⏰ Full Time 🟡 Mid-level 🟠 Senior 📊 Data Scientist Data Scientist III 🕒 2 days ago Conduent 10,000+ employees 🏥 Healthcare 📦 Logistics 💼 Consulting Website LinkedIn All Job Openings Technical lead delivering Azure-based Analytics, AI, and GenAI solutions for Conduent’s enterprise clients. Leading cloud-native development, RAG architectures, engineering standards, and developer mentorship in a hybrid India role. 🏢🏡 Noida – Hybrid 💰 Venture Round on 2009-01 ⏰ Full Time 🟠 Senior 🔴 Lead 📊 Data Scientist Data Scientist III 🕒 2 days ago Conduent 10,000+ employees 🏥 Healthcare 📦 Logistics 💼 Consulting Website LinkedIn All Job Openings Technical lead delivering Azure-based Analytics, AI, and GenAI solutions for Conduent’s mission-critical enterprise services. Designing RAG applications, cloud-native services, and reusable engineering platforms. 🏢🏡 Noida – Hybrid 💰 Venture Round on 2009-01 ⏰ Full Time 🟠 Senior 🔴 Lead 📊 Data Scientist View More Data Science Jobs 🌐 Worldwide Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or support@remoterocketship.com Search Search Jobs by country Search jobs by city Search jobs by job title Search entry-level jobs Search junior-level jobs Search senior-level jobs Search jobs by tech stack Search jobs by contract type Search remote internships Search remote part-time jobs Remote jobs Anywhere in the World Companies Hiring Anywhere in the World Companies Hiring Sales People Anywhere in the World Companies Hiring Software Engineers Anywhere in the World Resources Advice Tips for finding remote jobs Interview questions and answers Resume examples Cover letter examples Post a job Affiliates About us Is Remote Rocketship legit? Privacy policy Terms of service Job board SEO course Remote Job Search MasterClass Resume Review AI Apply Copilot OpenClaw job finder Find jobs using your resume Jobs by Country Remote jobs anywhere in the world (Worldwide remote jobs) Remote jobs United States Remote jobs Australia Remote jobs Brazil Remote jobs Canada Remote jobs France Remote jobs Ireland Remote jobs Germany Remote jobs Netherlands Remote jobs Spain Remote jobs UK Popular Jobs Remote data analyst jobs Remote customer support jobs Remote executive assistant jobs Remote marketing jobs Remote product designer jobs Remote product manager jobs Remote project manager jobs Remote recruiter jobs Remote sales jobs Remote software engineer jobs Jobs by Type Remote full-time jobs Remote part-time jobs Remote contract jobs Remote internship jobs Remote entry-level jobs Remote jobs with no experience required Remote junior jobs (1-3 years of experience) Digital nomad jobs Remote jobs with no degree required Freelance remote jobs Temporary remote jobs Remote jobs hiring now Stay at home mom jobs