Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Opedia Technologies

Inference Engineer

Opedia Technologies

Inference Engineer responsible for developing and optimizing AI model-serving systems. Collaborate with teams to ensure scalability and reliability for high-throughput, low-latency AI workloads.

Posted 7/21/2026full-timeBellevue • Washington • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production-grade machine learning inference systems, optimizing GPU utilization, and ensuring high throughput and low latency for AI workloads. Proficient in designing reliable distributed systems and implementing monitoring practices to maintain operational excellence.

Highest-signal resume keywords
Production Machine Learning Inference SystemsGPU Utilization OptimizationDistributed Systems DesignInference-Serving FrameworksKubernetes and Cloud Infrastructure

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Model-Serving SystemsPerformance Trade-Off AnalysisAPI-Based AI ProductsModel Optimization TechniquesLatency and Throughput Optimization
Soft Skills
Problem SolvingAdaptabilityCollaboration
Tools & Technologies
VLLMTensorRT-LLMTriton Inference ServerGPU SchedulingCloud Infrastructure Platforms
Industry Keywords
High-Volume Production ServicesLarge-Scale AI Serving PlatformsHyperscaler ExperienceAI Lab ExperienceML Infrastructure

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetes

About the role

Key responsibilities & impact
  • Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.
  • Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
  • Improve the scalability and operational maturity of inference platforms as customer demand grows.
  • Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.
  • Develop monitoring, alerting, and operational practices to maintain reliable inference services.
  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
  • Contribute to architecture decisions and engineering standards as the platform evolves.

Requirements

What you’ll need
  • Experience building and operating production machine learning inference or model-serving systems at scale.
  • Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.
  • Experience designing reliable distributed systems or production infrastructure.
  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.
  • Strong engineering fundamentals and the ability to independently own complex technical problems.
  • Comfortable working in a fast-moving environment where systems and processes are being built from the ground up.
  • Preferred Qualifications: Experience with modern inference-serving frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
  • Experience optimizing LLM inference workloads or large-scale AI serving platforms.
  • Background operating API-based AI products or high-volume production services.
  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.
  • Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning.
  • Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization.

Benefits

Comp & perks
  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
  • U.S. based employees have access to medical, dental, and vision insurance
  • 401(k) plan and company match
  • Employees also receive paid holidays, per calendar year.