FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Deep Learning Compiler Engineer – XLA
NVIDIADeep learning compiler engineer optimizing JAX and OpenXLA for NVIDIA GPUs. Building high-performance AI compiler software with framework and hardware teams.
Posted 8/14/2026full-timeRemote • California, Texas, Washington • 🇺🇸 United StatesSenior💰 $152,000 - $287,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing compiler optimization algorithms for deep learning workloads, with a strong focus on performance tuning and analysis for NVIDIA GPUs. Proficient in software engineering practices, mentoring, and collaboration with cross-functional teams.
Highest-signal resume keywords
Compiler Optimization AlgorithmsPerformance AnalysisC/C++ ProgrammingDeep Learning FrameworksNVIDIA GPU Programming
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Compiler OptimizationPerformance TuningSoftware EngineeringDebuggingTest DesignGraph PartitioningTensor ShardingCode GenerationDistributed ProgrammingDeep Learning Algorithms
Soft Skills
Interpersonal SkillsTeam CollaborationMentoring
Tools & Technologies
JAXOpenXLAMLIRLLVMOpenAI TritonCUDAOpenCLXLATVMPyTorch
Industry Keywords
High-Performance ComputingDeep LearningGPU ArchitectureAI Compiler Software
Tech Stack
Tools & technologiesPyTorchTensorflow
About the role
Key responsibilities & impact- Develop compiler optimization algorithms for deep learning workloads
- Optimize inference and training performance for the JAX framework and OpenXLA compiler on NVIDIA GPUs at scale
- Collaborate with deep learning framework partners and hardware architecture teams
- Craft and implement compiler optimization techniques for deep learning network graphs
- Design graph partitioning and tensor sharding techniques for distributed training and inference
- Perform performance tuning and analysis
- Implement code generation for NVIDIA GPU backends using MLIR, LLVM, and OpenAI Triton
- Design user-facing features in JAX and related libraries
- Perform general software engineering work
- Work with GPU hardware engineering teams to design AI compiler software features for next-generation GPUs
- Define project goals and scope and lead development efforts
- Apply software engineering and testing practices
- Mentor junior engineers and interns
Requirements
What you’ll need- Bachelor's, Master's, or Ph.D. in Computer Science, Computer Engineering, related field, or equivalent experience
- 4+ years of relevant work or research experience in performance analysis and compiler optimizations
- Ability to work independently, define project goals and scope, and lead development efforts
- Clean software engineering and testing practices
- Excellent C/C++ programming and software design skills
- Experience with debugging, performance analysis, and test design
- Strong foundation in CPU, GPU, or other high-performance hardware accelerator architecture
- Knowledge of high-performance computing and distributed programming
- CUDA or OpenCL programming experience desired but not required
- Experience with XLA, TVM, MLIR, LLVM, OpenAI Triton, deep learning models and algorithms, or deep learning framework design is a plus
- Strong interpersonal skills and ability to work in a dynamic product-oriented team
- Experience with JAX, PyTorch, or TensorFlow; CUDA or GPUs; and open-source compilers such as XLA, LLVM, MLIR, or TVM are ways to stand out
Benefits
Comp & perks- Competitive salaries
- Generous benefits package
- Equity
- Benefits