Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Opedia Technologies

AI Training Infrastructure Engineer

Opedia Technologies

Role involves building distributed systems for AI model training at a growing AI infrastructure company. Collaborate closely with engineering teams to solve complex challenges.

Posted 7/21/2026full-timeBellevue • Washington • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and scaling distributed training infrastructure for large AI models, with a strong focus on reliability, efficiency, and resource utilization. Proficient in integrating AI models into production pipelines and improving developer experience through automation and best practices.

Highest-signal resume keywords
Distributed Training SystemsLarge-Scale Machine Learning InfrastructureMulti-Node GPU TrainingProduction Machine Learning PipelinesProgramming Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed Training InfrastructureLarge AI ModelsFault ToleranceCheckpointingRecovery SolutionsTraining OperationsSystem ReliabilityScalability ChallengesEfficiency ImprovementComplex Distributed Systems
Soft Skills
OwnershipProblem-SolvingIndependence
Tools & Technologies
Automation ToolsAI Infrastructure PlatformPerformance Engineering
Industry Keywords
AI ModelsMachine LearningTraining PipelinesOperational Processes

Tech Stack

Tools & technologies
Distributed SystemsNode.js

About the role

Key responsibilities & impact
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
  • Design and improve systems that increase training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
  • Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.
  • Build tools and automation that improve the developer experience for AI researchers and engineers.
  • Establish best practices for training infrastructure, operational processes, and platform reliability.
  • Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.

Requirements

What you’ll need
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience working with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving engineering environment.
  • Comfortable operating with high ownership and limited process overhead.

Benefits

Comp & perks
  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
  • U.S. based employees have access to medical, dental, and vision insurance
  • 401(k) plan and company match
  • Paid holidays.