Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Inference Engineer

Designworks Talent LLC

Inference Engineer building and operating AI model-serving systems to support production-scale applications. Collaborating with engineering teams to optimize performance and reliability.

Posted 7/29/2026full-timeBellevue • Washington • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production-grade machine learning inference systems, optimizing for performance, scalability, and cost efficiency. Proficient in designing reliable distributed systems and managing GPU-backed AI workloads to ensure high throughput and low latency.

Highest-signal resume keywords
Machine Learning Inference SystemsGPU Utilization OptimizationDistributed Systems DesignPerformance Trade-offs AnalysisOperational Practices Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Model-Serving SystemsPerformance OptimizationScalability EngineeringReliability EngineeringMonitoring and AlertingInfrastructure DesignAI Workload ManagementTechnical Problem SolvingCost Efficiency AnalysisLatency and Throughput Management
Soft Skills
Independent OwnershipAdaptability in Fast-Moving Environments
Industry Keywords
AI WorkloadsInference InfrastructureProduction SystemsOperational MaturityArchitecture Decisions

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetes

About the role

Key responsibilities & impact
  • Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.
  • Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
  • Improve the scalability and operational maturity of inference platforms as customer demand grows.
  • Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.
  • Develop monitoring, alerting, and operational practices to maintain reliable inference services.
  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
  • Contribute to architecture decisions and engineering standards as the platform evolves.

Requirements

What you’ll need
  • Experience building and operating production machine learning inference or model-serving systems at scale.
  • Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.
  • Experience designing reliable distributed systems or production infrastructure.
  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.
  • Strong engineering fundamentals and the ability to independently own complex technical problems.
  • Comfortable working in a fast-moving environment where systems and processes are being built from the ground up.

Benefits

Comp & perks
  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.