FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer, Data Science – AI Accuracy
DroneDeploySenior Software Engineer focused on enhancing AI accuracy through data science and engineering. Building systems to maintain and optimize evaluation datasets in Auckland, New Zealand.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Python for data manipulation and analysis, SQL for database management, and a strong understanding of evaluation frameworks and prompt engineering. Capable of building robust datasets, conducting error analysis, and optimizing AI product performance through collaborative and iterative development.
Highest-signal resume keywords
Python Data ManipulationSQL Database ManagementEvaluation Dataset DevelopmentPrompt EngineeringError Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonSQLData ManipulationData AnalysisEvaluation FrameworksLabeling SystemsOntology DesignSchema DesignLarge DatasetsVisual Reasoning
Soft Skills
Collaborative MindsetIterative Development
Tools & Technologies
PandasNumPyBraintrustLangfuseAI-Assisted Tools
Industry Keywords
Data ScienceGround-Truth Quality ControlsMeasurement-Path IssuesAI ProductsIndustrial Inspection
Tech Stack
Tools & technologiesNumpyPandasPythonSQL
About the role
Key responsibilities & impact- Build and maintain robust dataset and evaluation infrastructure, including ground-truth quality controls
- Diagnose and fix measurement-path issues between offline evals and production accuracy
- Iterate on and scale existing labeling frameworks and knowledge capture systems
- Own prompt optimization across its full lifecycle, from candidate selection through production rollout and post-launch validation
- Automate eval pipelines and build the tooling to run this at scale across multiple AI products
- Create reusable internal tooling for error analysis, audits, and experimentation, and prioritize work using likely impact
Requirements
What you’ll need- Strong professional experience with Python for data manipulation, analysis, and tooling (e.g., pandas, NumPy)
- Solid SQL and database experience: querying, data modeling, and working with large datasets
- A data science background with a knowledge of LLM accuracy and in particular visual reasoning of these models
- Building and maintaining evaluation datasets, running and debugging evals with tools such as Braintrust or Langfuse, and conducting error analysis
- Experience designing ground-truth and labeling systems, and measuring label agreement and data quality
- Familiarity with prompt engineering and the full prompt lifecycle, from candidate selection through deployment and post-launch validation
- Ability to reason about ontology and schema design and its downstream impact on models, evals, and tooling
- Domain knowledge in construction or industrial inspection is a plus
- A collaborative mindset and a preference for iterative, team-oriented development
- Comfortable using AI-assisted tools for coding, data exploration, and debugging, while applying strong engineering judgment to guard against hallucinations and maintain high code quality.
Benefits
Comp & perks- Flexible schedules
- Family-friendly benefits
- Internal promotions