Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Perplexity

Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity

GPU infrastructure engineer building Perplexity's self-serve compute platform for AI training and inference. Operating multi-cloud Kubernetes clusters, scheduling scarce GPU capacity, and improving reliability.

Posted 8/5/2026full-timeSan Francisco • California • 🇺🇸 United StatesLead💰 $250,000 - $485,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and managing GPU cluster infrastructure, with a strong focus on Kubernetes, distributed systems, and high-availability services. Proficient in orchestrating compute across multiple cloud providers while ensuring workload resilience and observability.

Highest-signal resume keywords
Deep Kubernetes ExperienceGPU Cluster ManagementInfrastructure Code in Go, Rust, or C++Distributed Systems FundamentalsObservability for ML Workloads

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Kubernetes OperatorsCustom Resource Definitions (CRDs)NVIDIA Hardware KnowledgeCUDAHigh-Speed NetworkingMulti-Cloud OrchestrationScheduling and Resource AllocationFault ToleranceLong-Running Training JobsInference Serving Stacks
Soft Skills
End-to-End OwnershipProblem-Solving
Tools & Technologies
PrometheusGrafanaWeights & BiasesSlurmInfiniBandRoCERDMA
Industry Keywords
GPU Fleet ProvisioningCompute PlatformCloud ProvidersWorkload ResilienceCapacity Integration

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsGoGoogle Cloud PlatformGrafanaKubernetesPrometheusRust

About the role

Key responsibilities & impact
  • Build and own a self-serve compute platform for launching training jobs and operating inference services without direct GPU provisioning or cluster management
  • Operate GPU fleet provisioning, lifecycle management, reliability, and capacity integration across cloud providers
  • Build scheduling and placement logic to find and efficiently pack available GPU capacity and assign workloads to suitable hardware
  • Support long-running distributed training jobs and highly available, low-latency production inference services on the same fleet
  • Write Kubernetes operators and CRDs and manage multi-provider GPU clusters
  • Build fault tolerance, autoscaling, and observability for workload resilience and fleet utilization
  • Partner with inference and cloud infrastructure engineers on platform architecture and roadmap
  • Take end-to-end ownership of GPU cluster infrastructure at Perplexity, which serves hundreds of millions of queries monthly

Requirements

What you’ll need
  • Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation
  • Experience managing GPU clusters at scale
  • Knowledge of NVIDIA hardware, CUDA, and high-speed networking such as InfiniBand or RoCE
  • Experience orchestrating compute across multiple clouds, including CoreWeave, AWS, GCP, or similar
  • Strong distributed systems fundamentals in scheduling, resource allocation, and fault tolerance under load
  • Ability to write infrastructure and systems-level code in Go, Rust, or C++
  • Experience supporting long-running training jobs and high-availability inference services
  • Experience with inference serving stacks such as vLLM, SGLang, or TensorRT-LLM is valued
  • Experience with Slurm or other HPC schedulers is valued
  • GPU kernel experience in CUDA or Triton is valued but not required
  • Production experience with InfiniBand, RoCE, or RDMA is valued
  • Observability experience for ML workloads with Prometheus, Grafana, or Weights & Biases is valued
  • Ability to own problems end-to-end in an environment without a predetermined path

Benefits

Comp & perks
  • Equity
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Retirement benefits
  • Fitness benefits
  • Commuter benefits
  • Dependent care accounts
  • Regionally tailored benefits for full-time employees outside the U.S.