FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Engineer
MatterworksData Engineer building production pipelines that transform mass spectrometry data into reliable datasets. Supporting Matterworks’ AI models and scientific product for biological discovery.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing data pipelines, ensuring data quality, and collaborating with cross-functional teams. Proficient in Python and SQL, with a strong understanding of cloud data infrastructure and data handling best practices.
Highest-signal resume keywords
Data Pipeline DevelopmentPython ProficiencySQL ProficiencyCloud Data Infrastructure KnowledgeData Quality Assurance
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data Pipeline BuildingData Quality ChecksData IngestionData ReconciliationTesting and Monitoring
Soft Skills
Clear Written CommunicationCuriosity About Science
Tools & Technologies
Argo WorkflowsMetaflowEKSGlueAthenaApache IcebergParquetDuckDBTerraform
Industry Keywords
Life SciencesBiotechnologyBiochemistry
Tech Stack
Tools & technologiesApacheCloudPythonSQLTerraform
About the role
Key responsibilities & impact- Own well-scoped data pipelines end to end, including design, testing, instrumentation, documentation, and scheduled operation
- Turn raw mass spectrometry and molecular data into usable datasets with consistent schemas, trustworthy metadata, and documented definitions
- Design, build, and extend data quality checks that compare each build against the previous one before publication
- Stop pipelines when quality checks fail instead of shipping bad data
- Ingest new public and partner datasets by fetching, converting, validating, and reconciling them with existing data
- Improve systems continuously to handle messy scientific and vendor formats at scale
- Build datasets for AI training corpora and product or agentic-layer use
- Report to the Head of Engineering and collaborate daily with machine learning researchers, scientists, and the product team
- Grow from owning well-scoped pipelines and datasets toward owning larger parts of the data platform
Requirements
What you’ll need- 2+ years of professional experience building data pipelines in production
- Proficiency in Python and SQL
- Working knowledge of cloud data infrastructure
- Experience with or transferable knowledge of Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, and Terraform
- Experience owning a pipeline or dataset end to end, including tests, monitoring, and failure handling
- Comfort working with messy data and formats and reconciling disagreements between sources
- Daily use of AI coding tools, with skepticism about their output regarding production data correctness
- Clear written communication, especially when explaining failures and changes
- Curiosity about science
- Experience in life sciences, biotechnology, or biochemistry is a plus but not required
- Interest in contributing to an early-stage startup
- Candidates must currently be legally authorized to work in the United States
Benefits
Comp & perks- Stock options
- Health and dental benefits
- Vision benefits
- Long-term disability insurance
- Short-term disability insurance
- Life insurance
- 401k with company match
- Flexible work policy
- Unlimited time away policy
- Commuter benefits
- Parking
- Regular team meals and outings
- Company support for continued education/coursework
- Conference participation