Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Runpod

HPC Storage Engineer

Runpod

Senior Storage Engineer scaling distributed storage for Runpod’s AI developer cloud. Optimizing performance, reliability, automation, and networking for AI workloads.

Posted 9/9/2026full-timeRemote • 🇺🇸 United StatesSeniorLead💰 $180,000 - $260,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in infrastructure and storage engineering, with a strong focus on distributed storage systems, performance optimization, and production code development in Go or Python. Proven ability to lead capacity expansions and collaborate effectively across teams while maintaining high standards of observability and operational excellence.

Highest-signal resume keywords
Distributed Storage Systems ExpertiseProduction Code Development in Go or PythonPerformance Analysis and DebuggingObservability Tooling ExperienceLinux Internals and Storage-Stack Knowledge

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed Storage SystemsProduction Code DevelopmentPerformance AnalysisLinux InternalsStorage-Stack KnowledgeS3-Compatible Object StorageNetwork Tuning for StorageObservability ToolingCapacity ManagementInfrastructure as Code
Soft Skills
Self-StartingCollaborativeContinuous Improvement MindsetOwnership Across Team BoundariesLow-Ego with High Confidence
Tools & Technologies
PrometheusGrafanaDatadogKubernetesCSINVMeISCSIS3-Compatible APIsControl-Plane ServicesHigh-Speed IB/Ethernet Fabrics
Industry Keywords
Infrastructure EngineeringStorage EngineeringSystems EngineeringNetwork PerformanceData TieringPetabyte-Scale DeploymentCapacity ExpansionHardware RefreshMigrationsBlameless Post-Incident Follow-Through

Tech Stack

Tools & technologies
CloudGoGrafanaKubernetesLinuxNFSPrometheusPythonRust

About the role

Key responsibilities & impact
  • Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage
  • Tune device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client-side mount behavior
  • Diagnose complex performance problems end to end
  • Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption
  • Design and tune high-throughput storage network paths, including MTU, jumbo frames, congestion and flow control, multipath, and NIC/offload configuration
  • Optimize RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic
  • Collaborate with network engineering on topology, oversubscription, and cross-region data movement
  • Write production code in Go, Python, or similar for control-plane services, provisioning, data movement, and monitoring
  • Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs
  • Automate manual storage operations and manage infrastructure as code
  • Participate in code review, testing, and CI
  • Instrument the fleet for IOPS, throughput, latency, errors, retries, utilization, and per-tenant consumption
  • Build dashboards, SLOs, and alerts
  • Participate in storage on-call rotations and lead blameless post-incident follow-through
  • Help determine distributed storage systems, data tiering and placement, network tuning, and petabyte-scale purchasing and deployment

Requirements

What you’ll need
  • 8+ years in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale
  • Deep, practical experience with at least one distributed storage system — Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or comparable
  • Strong Linux internals and storage-stack knowledge: block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF
  • Experience building and/or operating S3-compatible object storage services
  • Solid networking fundamentals with specific experience tuning networks for storage workloads
  • Proficiency writing and shipping production code in Go, Python, Rust, or similar
  • Hands-on experience with observability tooling such as Prometheus, Grafana, or Datadog, including designing metrics
  • Track record of performance analysis and debugging under real production pressure
  • Self-starting with general direction
  • Continuous improvement mindset
  • Ownership across team boundaries
  • Collaborative and low-ego, with high confidence
  • Eligible to work in the United States
  • Unable to require employment visa sponsorship

Benefits

Comp & perks
  • Meaningful equity; everyone on the team receives stock options
  • Generous medical, dental & vision plans
  • Flexible PTO
  • Remote-first work environment
  • Slack-based internal communication
  • Passionate team on the cutting edge of AI infrastructure
  • $1,200 Home Office & Equipment Stipend