Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Solutions Architect – Large Scale AI Inference

NVIDIA

Senior Solutions Architect guiding EMEA customers deploying NVIDIA’s large-scale AI inference on GPU clusters. Optimizing MoE serving, interconnect-aware scheduling, and next-generation inference architectures.

Posted 9/8/2026full-timeRemote • 🇫🇷 FranceSenior💰 PLN 292,500 - PLN 507,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing large-scale AI inference workloads on multi-node GPU clusters, with a strong focus on transformers and MoE models. Engages effectively with technical teams and contributes to the AI developer community through workshops and collaborative projects.

Highest-signal resume keywords
Neural Networks Inference OptimizationTransformers Inference OptimizationMoE Inference At ScaleNVIDIA DynamoGPU Memory Hierarchies

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Inference Pipeline ArchitectureQuantizationSpeculative DecodingContinuous BatchingKV Cache OptimizationExpert ParallelismLoad BalancingAll-To-All CommunicationRouting OverheadHigh-Performance Computing
Soft Skills
Effective Engagement With ML EngineersCollaboration With Product Teams
Tools & Technologies
NVIDIA NIXLNVIDIA GroveDisaggregated Inference ToolingNVLinkInfiniBandRDMAUCX
Certifications & Qualifications
MS or PhD in Computer ScienceEngineeringHigh-Performance Computing
Industry Keywords
AI InferenceLarge-Scale AI InfrastructureTechnical WorkshopsHackathonsReference Architectures

Tech Stack

Tools & technologies
Node.js

About the role

Key responsibilities & impact
  • Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters
  • Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs
  • Improve inference efficiency across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments
  • Collaborate with NVIDIA product teams, including Dynamo, TensorRT-LLM, and NIXL, to accelerate customer success
  • Animate the AI inference developer community across EMEA through technical workshops, hackathons, and reference architectures
  • Establish technical direction for scalable, high-performance AI inference across demanding production environments

Requirements

What you’ll need
  • MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience
  • 5+ years of experience in Neural Networks inference optimization
  • Solid understanding of transformers inference optimization, including quantization, disaggregated inference, speculative decoding, continuous batching, and KV cache optimization
  • Practical experience in MoE inference at scale, including expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale
  • Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level
  • Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling
  • Understanding of GPU memory hierarchies and high-speed interconnects, including NVLink, InfiniBand, RDMA, and UCX
  • Contributions to advanced AI labs or large-scale AI infrastructure providers performing inference on thousands of GPUs
  • Published work or benchmarks in large-scale AI inference

Benefits

Comp & perks
  • Highly competitive salaries
  • Comprehensive benefits package