FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
PerplexityGPU infrastructure engineer building Perplexity's self-serve compute platform for AI training and inference. Operating multi-cloud Kubernetes clusters, scheduling scarce GPU capacity, and improving reliability.
Posted 8/5/2026full-timeSan Francisco • California • 🇺🇸 United StatesLead💰 $250,000 - $485,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing GPU cluster infrastructure, with a strong focus on Kubernetes, distributed systems, and high-availability services. Proficient in orchestrating compute across multiple cloud providers while ensuring workload resilience and observability.
Highest-signal resume keywords
Deep Kubernetes ExperienceGPU Cluster ManagementInfrastructure Code in Go, Rust, or C++Distributed Systems FundamentalsObservability for ML Workloads
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Kubernetes OperatorsCustom Resource Definitions (CRDs)NVIDIA Hardware KnowledgeCUDAHigh-Speed NetworkingMulti-Cloud OrchestrationScheduling and Resource AllocationFault ToleranceLong-Running Training JobsInference Serving Stacks
Soft Skills
End-to-End OwnershipProblem-Solving
Tools & Technologies
PrometheusGrafanaWeights & BiasesSlurmInfiniBandRoCERDMA
Industry Keywords
GPU Fleet ProvisioningCompute PlatformCloud ProvidersWorkload ResilienceCapacity Integration
Tech Stack
Tools & technologiesAWSCloudDistributed SystemsGoGoogle Cloud PlatformGrafanaKubernetesPrometheusRust
About the role
Key responsibilities & impact- Build and own a self-serve compute platform for launching training jobs and operating inference services without direct GPU provisioning or cluster management
- Operate GPU fleet provisioning, lifecycle management, reliability, and capacity integration across cloud providers
- Build scheduling and placement logic to find and efficiently pack available GPU capacity and assign workloads to suitable hardware
- Support long-running distributed training jobs and highly available, low-latency production inference services on the same fleet
- Write Kubernetes operators and CRDs and manage multi-provider GPU clusters
- Build fault tolerance, autoscaling, and observability for workload resilience and fleet utilization
- Partner with inference and cloud infrastructure engineers on platform architecture and roadmap
- Take end-to-end ownership of GPU cluster infrastructure at Perplexity, which serves hundreds of millions of queries monthly
Requirements
What you’ll need- Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation
- Experience managing GPU clusters at scale
- Knowledge of NVIDIA hardware, CUDA, and high-speed networking such as InfiniBand or RoCE
- Experience orchestrating compute across multiple clouds, including CoreWeave, AWS, GCP, or similar
- Strong distributed systems fundamentals in scheduling, resource allocation, and fault tolerance under load
- Ability to write infrastructure and systems-level code in Go, Rust, or C++
- Experience supporting long-running training jobs and high-availability inference services
- Experience with inference serving stacks such as vLLM, SGLang, or TensorRT-LLM is valued
- Experience with Slurm or other HPC schedulers is valued
- GPU kernel experience in CUDA or Triton is valued but not required
- Production experience with InfiniBand, RoCE, or RDMA is valued
- Observability experience for ML workloads with Prometheus, Grafana, or Weights & Biases is valued
- Ability to own problems end-to-end in an environment without a predetermined path
Benefits
Comp & perks- Equity
- Health insurance
- Dental insurance
- Vision insurance
- Retirement benefits
- Fitness benefits
- Commuter benefits
- Dependent care accounts
- Regionally tailored benefits for full-time employees outside the U.S.