Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
42dot

Senior AI Data Pipeline Engineer

42dot

Senior AI Data Pipeline Engineer creating scalable data pipelines for AI/ML initiatives. Collaborating with teams to optimize data processing workloads for mission-critical AI systems.

Posted 7/20/2026full-timePangyo • 🇰🇷 South KoreaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and building scalable data pipelines for AI and Machine Learning initiatives, with a strong focus on optimizing data processing using tools like Apache Spark and Databricks. Proficient in containerization with Kubernetes and workflow orchestration using Apache Airflow to ensure efficient data infrastructure.

Highest-signal resume keywords
Data Pipeline DevelopmentApache Spark ProficiencyKubernetes ContainerizationApache Airflow Workflow OrchestrationPython Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data Pipeline DesignDistributed Processing FrameworksCloud-Native ServicesData Infrastructure SecurityData Processing OptimizationEvent-Driven ArchitecturesComplex Dependency ManagementRoot Cause AnalysisData IngestionAI/ML Dataset Management
Soft Skills
Problem-SolvingCommunicationCollaboration
Tools & Technologies
DatabricksApache KafkaKubernetesApache AirflowSpark
Industry Keywords
AI InitiativesMachine LearningData InfrastructureScalable SystemsHigh-Throughput Data

Tech Stack

Tools & technologies
AirflowApacheCloudKafkaKubernetesPythonSpark

About the role

Key responsibilities & impact
  • Design and build high-performance, scalable data pipelines to support diverse AI and Machine Learning initiatives across the organization.
  • Architect and implement multi-region data infrastructure to ensure global data availability and seamless synchronization.
  • Develop flexible pipeline architectures that allow for complex branching and logic isolation to support multiple concurrent AI projects.
  • Optimize large-scale data processing workloads using Databricks and Spark to maximize throughput and minimize processing costs.
  • Maintain and evolve the containerized data environment on Kubernetes, ensuring robust and reliable execution of data workloads.
  • Collaborate with AI researchers and platform teams to streamline the flow of high-quality data into training and evaluation pipelines.

Requirements

What you’ll need
  • Extensive professional experience in building and operating production-grade data pipelines for massive-scale AI/ML datasets.
  • Strong proficiency in distributed processing frameworks, particularly Apache Spark and the Databricks ecosystem.
  • Deep hands-on experience with workflow orchestration tools like Apache Airflow for managing complex dependency graphs.
  • Solid understanding of Kubernetes and containerization for deploying and scaling data processing components.
  • Proficiency in distributed messaging systems such as Apache Kafka for high-throughput data ingestion and event-driven architectures.
  • Expert-level programming skills in Python for system-level optimizations.
  • Strong knowledge of cloud-native services and best practices for building secure and scalable data infrastructure.
  • Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems.
  • Strong communication skills to effectively collaborate with cross-functional teams and external partners.

Benefits

Comp & perks
  • Health insurance
  • 401(k) matching
  • Flexible work hours
  • Paid time off
  • Professional development opportunities