FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Inference Engineer
Designworks Talent LLCInference Engineer building and operating AI model-serving systems to support production-scale applications. Collaborating with engineering teams to optimize performance and reliability.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production-grade machine learning inference systems, optimizing for performance, scalability, and cost efficiency. Proficient in designing reliable distributed systems and managing GPU-backed AI workloads to ensure high throughput and low latency.
Highest-signal resume keywords
Machine Learning Inference SystemsGPU Utilization OptimizationDistributed Systems DesignPerformance Trade-offs AnalysisOperational Practices Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model-Serving SystemsPerformance OptimizationScalability EngineeringReliability EngineeringMonitoring and AlertingInfrastructure DesignAI Workload ManagementTechnical Problem SolvingCost Efficiency AnalysisLatency and Throughput Management
Soft Skills
Independent OwnershipAdaptability in Fast-Moving Environments
Industry Keywords
AI WorkloadsInference InfrastructureProduction SystemsOperational MaturityArchitecture Decisions
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetes
About the role
Key responsibilities & impact- Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.
- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.
- Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
- Improve the scalability and operational maturity of inference platforms as customer demand grows.
- Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.
- Develop monitoring, alerting, and operational practices to maintain reliable inference services.
- Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
- Contribute to architecture decisions and engineering standards as the platform evolves.
Requirements
What you’ll need- Experience building and operating production machine learning inference or model-serving systems at scale.
- Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.
- Experience designing reliable distributed systems or production infrastructure.
- Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.
- Strong engineering fundamentals and the ability to independently own complex technical problems.
- Comfortable working in a fast-moving environment where systems and processes are being built from the ground up.
Benefits
Comp & perks- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.