Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Fuse Energy

AI Inference Engineer

Fuse Energy

Founding AI Inference Engineer at Fuse Energy defining and building AI inference serving architecture. Collaborating with CUDA/GPU engineers at a renewable energy startup.

Posted 7/20/2026full-timeRemote • 🇬🇧 United KingdomMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and building large-scale inference serving systems, with a focus on performance optimization techniques such as quantisation and batching. Proven ability to collaborate with GPU/CUDA engineers and make critical architecture decisions in a dynamic, early-stage environment.

Highest-signal resume keywords
Inference Serving FrameworksPerformance Optimization TechniquesGPU/CUDA IntegrationArchitecture Decision-MakingCapacity Planning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
BatchingKV-Cache ManagementQuantisationSpeculative DecodingAutoscalingTriton Inference ServerTensorRT-LLMVLLMSGLangLarge-Scale Inference Systems
Soft Skills
Systems ThinkingProblem SolvingAdaptability
Tools & Technologies
KubernetesSlurm
Industry Keywords
Multi-Tenant ServingSLA-Driven InfrastructureEnergy MarketsGrid SystemsSustainability-Focused Compute

Tech Stack

Tools & technologies
Kubernetes

About the role

Key responsibilities & impact
  • Define Fuse's inference serving strategy and architecture from first principles.
  • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
  • Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
  • Set the standards, tooling, and benchmarks this function will run on as it grows.

Requirements

What you’ll need
  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • A track record of making high-stakes architecture calls and owning the outcome.
  • Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.
  • **Nice to Have**
  • Experience with Triton or custom ML inference/training frameworks.
  • Experience with autoscaling or capacity planning for large-scale inference workloads.
  • Exposure to multi-tenant serving or SLA-driven infrastructure.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
  • Familiarity with Kubernetes/Slurm for cluster orchestration.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute.

Benefits

Comp & perks
  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office based employees.