FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Member of Technical Staff – AI Cloud Infrastructure
Emerald AISenior infrastructure engineer architecting Emerald AI’s flexible managed cloud for AI data centers. Building Kubernetes, Slurm, multi-tenant GPU, storage, and billing platforms.
Posted 8/6/2026full-timeSan Francisco • California, District of Columbia, Massachusetts • 🇺🇸 United StatesLeadWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and launching managed cloud platforms, with a strong focus on Kubernetes, Slurm, and parallel filesystem solutions. Proficient in infrastructure-as-code practices and capable of integrating GPU infrastructure for high-performance computing environments.
Highest-signal resume keywords
Kubernetes ManagementSlurm AdministrationParallel Filesystem DeploymentInfrastructure-as-Code with TerraformPython or Go Programming
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Infrastructure EngineeringCloud Service FundamentalsMulti-Tenancy DesignLinux Systems KnowledgeGPU Infrastructure FamiliarityControl Plane ArchitectureAPIs and Metering SystemsOperational Service DisciplineIncident Response ProtocolsUsage Metering
Tools & Technologies
TerraformAnsibleLustreVASTWekaInfiniBandRoCERDMAS3Ceph
Industry Keywords
Managed CloudAI PlatformHigh-Performance ComputingGPU CloudHyperscaler AI Service
Tech Stack
Tools & technologiesAnsibleCloudGoKubernetesLinuxNode.jsPythonTerraform
About the role
Key responsibilities & impact- Architect managed services from 0→1, including GPU capacity productization, isolation boundaries, tenant models, provisioning flows, and service catalogs
- Build control-plane services, self-service customer interfaces, automated lifecycle systems, and usage metering integrated with billing
- Onboard and technically assess bare-metal GPU infrastructure partners
- Design and implement multi-tenancy across compute, storage, and networking with secure isolation, QoS, and encryption
- Manage Kubernetes and Slurm environments for large-scale training and inference
- Oversee node health, driver fleets, and kernel management across heterogeneous clouds
- Deploy and integrate parallel storage solutions such as Lustre, VAST, or Weka
- Define SLOs, observability standards, and incident response protocols
- Bridge internal operational standards with provider SLAs to deliver a reliable managed cloud product
Requirements
What you’ll need- At least 7+ years of infrastructure or platform engineering experience
- Experience architecting and launching a managed cloud or AI platform that reached production users
- Strong experience with Kubernetes and Slurm as managed services
- Production experience deploying or operating Lustre or comparable parallel filesystems such as GPFS, Weka, VAST, or BeeGFS
- Understanding of parallel filesystem architecture, tuning, and failure modes
- Strong grasp of cloud service fundamentals, control planes, tenancy and isolation models, APIs, quota and metering systems, and operational service discipline
- Deep Linux systems knowledge
- Mature infrastructure-as-code practice with Terraform and Ansible
- Programming ability in Python or Go
- Familiarity with GPU infrastructure and high-performance networking using InfiniBand, RoCE, and RDMA
- Familiarity with the GPU software stack
- Preferred: experience at a GPU cloud, hyperscaler AI service, or HPC center delivering compute and storage as a service
- Preferred: familiarity with NVIDIA SuperPOD, GPUDirect Storage, NCCL debugging, and DCGM
- Preferred: experience with Lustre multitenancy features or service-provider VAST/Weka deployments
- Preferred: experience negotiating with and integrating multiple infrastructure vendors
- Preferred: experience running object storage at scale with S3, Ceph, or MinIO, including data tiering
- Preferred: experience building billing, metering, or FinOps pipelines
- Must be located in or willing to locate to an Emerald AI hub
- Must address ability to work in the United States
Benefits
Comp & perks- Competitive pay + equity
- Stock options
- Medical benefits
- Dental benefits
- Vision benefits
- 401(k) matching
- Flexible location
- 2 WFH days/week
- Collaborative, low-ego environment
- Opportunity to influence strategy, GTM, org design, and customer/investor engagement from day one
- Reasonable accommodations for applicants with disabilities