Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Baseten

Engineering

Baseten

Baseten Embeddings Inference: The fastest embeddings solution available; The Baseten Inference Stack

Posted 7/9/2026Verified active Jul 25, 2026, 12:37 AMfull-timeSan Francisco • California • 🇺🇸 United StatesWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Candidates should emphasize their leadership experience in managing GPU or ML systems engineering teams, along with a strong technical background in GPU kernel engineering and CUDA development. Highlighting the ability to set technical direction, optimize performance, and communicate effectively across teams will be crucial.

Highest-signal resume keywords
Leadership in GPU or ML Systems EngineeringTechnical Direction SettingCUDA Kernel DevelopmentPerformance OptimizationTeam Development and Hiring

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDAGPU Kernel EngineeringGEMMsAttention MechanismsMoE RoutingProfiling MethodologyQuantization (FP8/FP4)TritonCUTLASSCuTe DSL
Soft Skills
Team LeadershipTechnical CommunicationRoadmap PrioritizationMentoringProcess Building
Tools & Technologies
NVIDIA GPU ArchitecturesCUDA EcosystemBaseten Inference StackOpen-source GPU LibrariesInference Frameworks
Industry Keywords
GPU ArchitectureMachine Learning SystemsProduction InferenceLatency ReductionCost Optimization

Tech Stack

Tools & technologies
Spark

About the role

Key responsibilities & impact
  • We're looking for an Engineering Manager to lead our GPU Kernel Engineering team, the group responsible for writing the low-level CUDA code that makes Baseten's inference stack faster than anyone else's. This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers.
  • You'll own the technical direction of a team working at the intersection of GPU architecture, ML systems, and production inference. Your engineers write CUDA kernels for GEMMs, attention mechanisms, and MoE routing, optimize at the warp and tensor-core level, and ship improvements that directly reduce latency and cost for the AI companies running their most critical workloads on Baseten.
  • This role is not for someone who wants to step away from the technical work. You'll be close enough to the code to credibly review it, set direction, and unblock your team, while also building the processes, culture, and roadmap that let a world-class kernel team operate at its best.

Requirements

What you’ll need
  • Proven experience leading a team of GPU or ML systems engineers, with a track record of hiring and developing strong technical talent
  • Deep personal background in GPU kernel engineering. You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level
  • Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology
  • Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem
  • Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership
  • Hands-on experience with Triton, CUTLASS, or CuTe DSL
  • Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing
  • Open-source contributions to GPU libraries or inference frameworks
  • Experience presenting technical work at NVIDIA GTC, MLSys, or similar venues

Benefits

Comp & perks
  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
  • At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
  • We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).