FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Solutions Architect – Large Scale AI Inference
NVIDIASenior Solutions Architect guiding EMEA customers deploying NVIDIA’s large-scale AI inference on GPU clusters. Optimizing MoE serving, interconnect-aware scheduling, and next-generation inference architectures.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in optimizing large-scale AI inference workloads on multi-node GPU clusters, with a strong focus on transformers and MoE models. Engages effectively with technical teams and contributes to the AI developer community through workshops and collaborative projects.
Highest-signal resume keywords
Neural Networks Inference OptimizationTransformers Inference OptimizationMoE Inference At ScaleNVIDIA DynamoGPU Memory Hierarchies
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Inference Pipeline ArchitectureQuantizationSpeculative DecodingContinuous BatchingKV Cache OptimizationExpert ParallelismLoad BalancingAll-To-All CommunicationRouting OverheadHigh-Performance Computing
Soft Skills
Effective Engagement With ML EngineersCollaboration With Product Teams
Tools & Technologies
NVIDIA NIXLNVIDIA GroveDisaggregated Inference ToolingNVLinkInfiniBandRDMAUCX
Certifications & Qualifications
MS or PhD in Computer ScienceEngineeringHigh-Performance Computing
Industry Keywords
AI InferenceLarge-Scale AI InfrastructureTechnical WorkshopsHackathonsReference Architectures
Tech Stack
Tools & technologiesNode.js
About the role
Key responsibilities & impact- Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters
- Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs
- Improve inference efficiency across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments
- Collaborate with NVIDIA product teams, including Dynamo, TensorRT-LLM, and NIXL, to accelerate customer success
- Animate the AI inference developer community across EMEA through technical workshops, hackathons, and reference architectures
- Establish technical direction for scalable, high-performance AI inference across demanding production environments
Requirements
What you’ll need- MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience
- 5+ years of experience in Neural Networks inference optimization
- Solid understanding of transformers inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache optimization
- Practical experience in MoE inference at scale, including expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale
- Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level
- Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling
- Understanding of GPU memory hierarchies and high-speed interconnects, including NVLink, InfiniBand, RDMA, and UCX
- Contributions to advanced AI labs or large-scale AI infrastructure providers performing inference on thousands of GPUs
- Published work or benchmarks in large-scale AI inference
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package