Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Superintelligence

Member of Technical Staff, Performance & Capacity

Superintelligence

Performance and capacity engineer optimizing GPU fleets, networking, storage, and workloads. Supporting Physical Superintelligence’s AI systems for scalable scientific discovery and new physics.

Posted 8/16/2026full-timeBoston • Massachusetts • 🇺🇸 United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing GPU performance for large-scale compute workloads, with a strong focus on capacity modeling, parallel file systems, and interconnect behavior. Capable of clearly communicating complex performance results to researchers and making informed capacity decisions with financial implications.

Highest-signal resume keywords
GPU Performance OptimizationDistributed Training PerformanceCapacity Decision OwnershipParallel File Systems KnowledgeInfiniBand Networking Expertise

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GPU WorkloadsLarge-Scale ComputePerformance Bottleneck AnalysisAI Training and InferenceCommunication/Compute OverlapBatchingLatency FloorsError Bar AnalysisCheckpointing EconomicsData Locality Optimization
Soft Skills
Clear CommunicationProblem-Solving
Tools & Technologies
InfiniBandRDMA NetworkingParallel File Systems
Industry Keywords
Capacity ModelingMulti-Node GPU TrainingTraining ThroughputHPC BackgroundData Center Experience

Tech Stack

Tools & technologies
CloudNode.js

About the role

Key responsibilities & impact
  • Measure fleet utilization, goodput, and training efficiency on real workloads
  • Identify and close sources of wasted capacity
  • Own the capacity model, including node counts, assumptions, and error bars
  • Defend capacity decisions when real money is committed
  • Own workload mix, preemptible fraction, and checkpointing economics over time
  • Optimize parallel file systems, data locality, and interconnect behavior for multi-node GPU training
  • Locate actual performance bottlenecks beyond profiler summaries
  • Explain performance results to researchers in usable terms

Requirements

What you’ll need
  • Five or more years with GPU and large-scale compute workloads
  • Real multi-node experience with distributed training performance, interconnects, and collective communication
  • System-level knowledge of GPU performance characteristics, including H100 and B200 workloads
  • Deep understanding of AI training and inference performance
  • Understanding of step time, MFU, communication/compute overlap, decode bandwidth, batching, and latency floors
  • Experience owning a capacity decision with financial consequences
  • Fluent knowledge of parallel file systems and their impact on training throughput
  • Fluent knowledge of InfiniBand or RDMA-class networking behavior
  • Ability to explain performance results clearly to researchers
  • Work authorization/sponsorship may be relevant for the United States
  • Nice to have: data center floor experience, thermal and power envelopes, failure domains, energy- or preemption-aware scheduling, cloud and specialist GPU provider sourcing, and scientific computing or HPC background

Benefits

Comp & perks
  • Meaningful early-stage equity
  • Competitive compensation
  • Benefits
  • Full ownership of work from spec to ship to on-call
  • AI-native development process with daily use of agentic coding tools