Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Deep Learning Compiler Engineer – XLA

NVIDIA

Deep learning compiler engineer optimizing JAX and OpenXLA for NVIDIA GPUs. Building high-performance AI compiler software with framework and hardware teams.

Posted 8/14/2026full-timeRemote • California, Texas, Washington • 🇺🇸 United StatesSenior💰 $152,000 - $287,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing compiler optimization algorithms for deep learning workloads, with a strong focus on performance tuning and analysis for NVIDIA GPUs. Proficient in software engineering practices, mentoring, and collaboration with cross-functional teams.

Highest-signal resume keywords
Compiler Optimization AlgorithmsPerformance AnalysisC/C++ ProgrammingDeep Learning FrameworksNVIDIA GPU Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Compiler OptimizationPerformance TuningSoftware EngineeringDebuggingTest DesignGraph PartitioningTensor ShardingCode GenerationDistributed ProgrammingDeep Learning Algorithms
Soft Skills
Interpersonal SkillsTeam CollaborationMentoring
Tools & Technologies
JAXOpenXLAMLIRLLVMOpenAI TritonCUDAOpenCLXLATVMPyTorch
Industry Keywords
High-Performance ComputingDeep LearningGPU ArchitectureAI Compiler Software

Tech Stack

Tools & technologies
PyTorchTensorflow

About the role

Key responsibilities & impact
  • Develop compiler optimization algorithms for deep learning workloads
  • Optimize inference and training performance for the JAX framework and OpenXLA compiler on NVIDIA GPUs at scale
  • Collaborate with deep learning framework partners and hardware architecture teams
  • Craft and implement compiler optimization techniques for deep learning network graphs
  • Design graph partitioning and tensor sharding techniques for distributed training and inference
  • Perform performance tuning and analysis
  • Implement code generation for NVIDIA GPU backends using MLIR, LLVM, and OpenAI Triton
  • Design user-facing features in JAX and related libraries
  • Perform general software engineering work
  • Work with GPU hardware engineering teams to design AI compiler software features for next-generation GPUs
  • Define project goals and scope and lead development efforts
  • Apply software engineering and testing practices
  • Mentor junior engineers and interns

Requirements

What you’ll need
  • Bachelor's, Master's, or Ph.D. in Computer Science, Computer Engineering, related field, or equivalent experience
  • 4+ years of relevant work or research experience in performance analysis and compiler optimizations
  • Ability to work independently, define project goals and scope, and lead development efforts
  • Clean software engineering and testing practices
  • Excellent C/C++ programming and software design skills
  • Experience with debugging, performance analysis, and test design
  • Strong foundation in CPU, GPU, or other high-performance hardware accelerator architecture
  • Knowledge of high-performance computing and distributed programming
  • CUDA or OpenCL programming experience desired but not required
  • Experience with XLA, TVM, MLIR, LLVM, OpenAI Triton, deep learning models and algorithms, or deep learning framework design is a plus
  • Strong interpersonal skills and ability to work in a dynamic product-oriented team
  • Experience with JAX, PyTorch, or TensorFlow; CUDA or GPUs; and open-source compilers such as XLA, LLVM, MLIR, or TVM are ways to stand out

Benefits

Comp & perks
  • Competitive salaries
  • Generous benefits package
  • Equity
  • Benefits