FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Inference Engineer
Fuse EnergyFounding AI Inference Engineer at Fuse Energy defining and building AI inference serving architecture. Collaborating with CUDA/GPU engineers at a renewable energy startup.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and building large-scale inference serving systems, with a focus on performance optimization techniques such as quantisation and batching. Proven ability to collaborate with GPU/CUDA engineers and make critical architecture decisions in a dynamic, early-stage environment.
Highest-signal resume keywords
Inference Serving FrameworksPerformance Optimization TechniquesGPU/CUDA IntegrationArchitecture Decision-MakingCapacity Planning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
BatchingKV-Cache ManagementQuantisationSpeculative DecodingAutoscalingTriton Inference ServerTensorRT-LLMVLLMSGLangLarge-Scale Inference Systems
Soft Skills
Systems ThinkingProblem SolvingAdaptability
Tools & Technologies
KubernetesSlurm
Industry Keywords
Multi-Tenant ServingSLA-Driven InfrastructureEnergy MarketsGrid SystemsSustainability-Focused Compute
Tech Stack
Tools & technologiesKubernetes
About the role
Key responsibilities & impact- Define Fuse's inference serving strategy and architecture from first principles.
- Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
- Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
- Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
- Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
- Act as a direct technical owner of inference performance and reliability.
- Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
- Set the standards, tooling, and benchmarks this function will run on as it grows.
Requirements
What you’ll need- 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
- Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
- Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
- Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
- A track record of making high-stakes architecture calls and owning the outcome.
- Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.
- **Nice to Have**
- Experience with Triton or custom ML inference/training frameworks.
- Experience with autoscaling or capacity planning for large-scale inference workloads.
- Exposure to multi-tenant serving or SLA-driven infrastructure.
- Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
- Familiarity with Kubernetes/Slurm for cluster orchestration.
- Interest or experience in energy markets, grid systems, or sustainability-focused compute.
Benefits
Comp & perks- Competitive salary and an equity sign-on bonus.
- Biannual bonus scheme.
- Fully expensed tech to match your needs.
- Breakfast and dinner allowance for office based employees.