Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Emerald AI

Member of Technical Staff – AI Cloud Infrastructure

Emerald AI

Senior infrastructure engineer architecting Emerald AI’s flexible managed cloud for AI data centers. Building Kubernetes, Slurm, multi-tenant GPU, storage, and billing platforms.

Posted 8/6/2026full-timeSan Francisco • California, District of Columbia, Massachusetts • 🇺🇸 United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting and launching managed cloud platforms, with a strong focus on Kubernetes, Slurm, and parallel filesystem solutions. Proficient in infrastructure-as-code practices and capable of integrating GPU infrastructure for high-performance computing environments.

Highest-signal resume keywords
Kubernetes ManagementSlurm AdministrationParallel Filesystem DeploymentInfrastructure-as-Code with TerraformPython or Go Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure EngineeringCloud Service FundamentalsMulti-Tenancy DesignLinux Systems KnowledgeGPU Infrastructure FamiliarityControl Plane ArchitectureAPIs and Metering SystemsOperational Service DisciplineIncident Response ProtocolsUsage Metering
Tools & Technologies
TerraformAnsibleLustreVASTWekaInfiniBandRoCERDMAS3Ceph
Industry Keywords
Managed CloudAI PlatformHigh-Performance ComputingGPU CloudHyperscaler AI Service

Tech Stack

Tools & technologies
AnsibleCloudGoKubernetesLinuxNode.jsPythonTerraform

About the role

Key responsibilities & impact
  • Architect managed services from 0→1, including GPU capacity productization, isolation boundaries, tenant models, provisioning flows, and service catalogs
  • Build control-plane services, self-service customer interfaces, automated lifecycle systems, and usage metering integrated with billing
  • Onboard and technically assess bare-metal GPU infrastructure partners
  • Design and implement multi-tenancy across compute, storage, and networking with secure isolation, QoS, and encryption
  • Manage Kubernetes and Slurm environments for large-scale training and inference
  • Oversee node health, driver fleets, and kernel management across heterogeneous clouds
  • Deploy and integrate parallel storage solutions such as Lustre, VAST, or Weka
  • Define SLOs, observability standards, and incident response protocols
  • Bridge internal operational standards with provider SLAs to deliver a reliable managed cloud product

Requirements

What you’ll need
  • At least 7+ years of infrastructure or platform engineering experience
  • Experience architecting and launching a managed cloud or AI platform that reached production users
  • Strong experience with Kubernetes and Slurm as managed services
  • Production experience deploying or operating Lustre or comparable parallel filesystems such as GPFS, Weka, VAST, or BeeGFS
  • Understanding of parallel filesystem architecture, tuning, and failure modes
  • Strong grasp of cloud service fundamentals, control planes, tenancy and isolation models, APIs, quota and metering systems, and operational service discipline
  • Deep Linux systems knowledge
  • Mature infrastructure-as-code practice with Terraform and Ansible
  • Programming ability in Python or Go
  • Familiarity with GPU infrastructure and high-performance networking using InfiniBand, RoCE, and RDMA
  • Familiarity with the GPU software stack
  • Preferred: experience at a GPU cloud, hyperscaler AI service, or HPC center delivering compute and storage as a service
  • Preferred: familiarity with NVIDIA SuperPOD, GPUDirect Storage, NCCL debugging, and DCGM
  • Preferred: experience with Lustre multitenancy features or service-provider VAST/Weka deployments
  • Preferred: experience negotiating with and integrating multiple infrastructure vendors
  • Preferred: experience running object storage at scale with S3, Ceph, or MinIO, including data tiering
  • Preferred: experience building billing, metering, or FinOps pipelines
  • Must be located in or willing to locate to an Emerald AI hub
  • Must address ability to work in the United States

Benefits

Comp & perks
  • Competitive pay + equity
  • Stock options
  • Medical benefits
  • Dental benefits
  • Vision benefits
  • 401(k) matching
  • Flexible location
  • 2 WFH days/week
  • Collaborative, low-ego environment
  • Opportunity to influence strategy, GTM, org design, and customer/investor engagement from day one
  • Reasonable accommodations for applicants with disabilities