Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Elorian AI

Inference Infrastructure Engineer, Serving

Elorian AI

Infrastructure Engineer design and optimize systems serving large multimodal AI models at an early-stage AI lab. Focus on improving performance to facilitate research and deployment.

Posted 7/27/2026full-timePalo Alto • California • 🇺🇸 United StatesMid-LevelSenior💰 $200,000 - $400,000 per yearWebsite

Tech Stack

Tools & technologies
Node.jsPython

About the role

Key responsibilities & impact
  • Build low-latency, high-throughput inference serving systems for our large multimodal models
  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management
  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)
  • Build autoscaling and load balancing for production ML services
  • Establish standards for reliability, observability, and reproducibility across the inference stack
  • Collaborate with researchers to enable high-performance inference for novel architectures

Requirements

What you’ll need
  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)
  • Strong systems programming skills; C++/CUDA a plus alongside Python
  • Experience with autoscaling and load balancing for production ML services
  • A track record of GPU cost optimization at scale
  • Preferred qualifications (strong candidates may have some, not all): Experience serving multimodal (vision + language) models
  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)

Benefits

Comp & perks
  • health, dental, and vision benefits
  • unlimited PTO
  • paid parental leave
  • relocation support as needed