FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Member of Technical Staff, Performance & Capacity
SuperintelligencePerformance and capacity engineer optimizing GPU fleets, networking, storage, and workloads. Supporting Physical Superintelligence’s AI systems for scalable scientific discovery and new physics.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in optimizing GPU performance for large-scale compute workloads, with a strong focus on capacity modeling, parallel file systems, and interconnect behavior. Capable of clearly communicating complex performance results to researchers and making informed capacity decisions with financial implications.
Highest-signal resume keywords
GPU Performance OptimizationDistributed Training PerformanceCapacity Decision OwnershipParallel File Systems KnowledgeInfiniBand Networking Expertise
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GPU WorkloadsLarge-Scale ComputePerformance Bottleneck AnalysisAI Training and InferenceCommunication/Compute OverlapBatchingLatency FloorsError Bar AnalysisCheckpointing EconomicsData Locality Optimization
Soft Skills
Clear CommunicationProblem-Solving
Tools & Technologies
InfiniBandRDMA NetworkingParallel File Systems
Industry Keywords
Capacity ModelingMulti-Node GPU TrainingTraining ThroughputHPC BackgroundData Center Experience
Tech Stack
Tools & technologiesCloudNode.js
About the role
Key responsibilities & impact- Measure fleet utilization, goodput, and training efficiency on real workloads
- Identify and close sources of wasted capacity
- Own the capacity model, including node counts, assumptions, and error bars
- Defend capacity decisions when real money is committed
- Own workload mix, preemptible fraction, and checkpointing economics over time
- Optimize parallel file systems, data locality, and interconnect behavior for multi-node GPU training
- Locate actual performance bottlenecks beyond profiler summaries
- Explain performance results to researchers in usable terms
Requirements
What you’ll need- Five or more years with GPU and large-scale compute workloads
- Real multi-node experience with distributed training performance, interconnects, and collective communication
- System-level knowledge of GPU performance characteristics, including H100 and B200 workloads
- Deep understanding of AI training and inference performance
- Understanding of step time, MFU, communication/compute overlap, decode bandwidth, batching, and latency floors
- Experience owning a capacity decision with financial consequences
- Fluent knowledge of parallel file systems and their impact on training throughput
- Fluent knowledge of InfiniBand or RDMA-class networking behavior
- Ability to explain performance results clearly to researchers
- Work authorization/sponsorship may be relevant for the United States
- Nice to have: data center floor experience, thermal and power envelopes, failure domains, energy- or preemption-aware scheduling, cloud and specialist GPU provider sourcing, and scientific computing or HPC background
Benefits
Comp & perks- Meaningful early-stage equity
- Competitive compensation
- Benefits
- Full ownership of work from spec to ship to on-call
- AI-native development process with daily use of agentic coding tools