FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Pipeline Engineer
PeerIslandsData Pipeline Engineer building and operating data ingestion systems for a large-scale benefits administration platform. Collaborating on data quality, identity resolution, and transformation processes while ensuring operational excellence.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production data pipelines at scale, with a strong focus on data quality frameworks, entity resolution, and data mapping. Proficient in utilizing technologies such as Kafka, SQL, and programming languages like Java and Python for effective data ingestion and processing.
Highest-signal resume keywords
Production Data Pipeline DevelopmentKafka Consumers/ProducersSQL ProficiencyData Quality FrameworksEntity Resolution/Master Data Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data Pipeline DevelopmentSQLJavaPythonETL FundamentalsMedallion LayeringData Quality GatesEntity ResolutionData MappingCrosswalk Discipline
Tools & Technologies
KafkaInformatica MDMReltioYAMLJSONGit
Industry Keywords
Batch-Stream Hand-OffWatermark PatternsCDCDeduplicationSurvivorship Rules
Tech Stack
Tools & technologiesETLInformaticaJavaKafkaPythonSQL
About the role
Key responsibilities & impact- Build batch-seed and event-tail ingestion per source system, including seed→tail watermark hand-off, idempotent upserts, and dedup ledgers
- Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints, retry/backoff, DLQ), and full observability
- Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation — financial is zero-tolerance
- Build identity matching combining deterministic rules with probabilistic scoring and confidence bands; deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data
- Author and maintain source→canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration
- Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution
Requirements
What you’ll need- 5+ years building production data pipelines at scale
- Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns
- Strong SQL and solid ETL fundamentals
- Java and/or Python in production
- Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, and batch–stream hand-off
- Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation
- Entity resolution / MDM exposure: record matching, dedup, survivorship — via commercial tools (Informatica MDM, Reltio) or custom builds
- Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data, config-as-code (YAML/JSON, Git)
Benefits
Comp & perks- Health insurance
- Retirement plans
- Paid time off
- Flexible work arrangements
- Professional development