FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Cluster Engineer
STN IncorporatedAI infrastructure engineer designing and optimizing large-scale GPU clusters for AI training and inference. Tuning distributed frameworks, networking, storage, scheduling, and performance across production GPU environments.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive expertise in designing and optimizing multi-node GPU clusters for AI training and inference, with a strong focus on performance tuning, benchmarking, and resource allocation. Proficient in distributed training environments and Linux infrastructure management, ensuring high efficiency and scalability.
Highest-signal resume keywords
Multi-Node GPU Cluster DesignDistributed PyTorch TrainingNCCL Communication OptimizationLinux Systems AdministrationGPU Inference Throughput Optimization
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GPU Utilization OptimizationCUDA TuningNCCL BenchmarkingPython ScriptingBash ScriptingAI Storage DesignPerformance Bottleneck AnalysisData ParallelismNetwork Topology OptimizationHigh-Performance Parallel I/O
Tools & Technologies
SlurmPyTorchNVIDIA DCGMNsight SystemsPyxisEnrootKubernetesAWSAzureGCP
Industry Keywords
AI InfrastructureHPC WorkloadsGPU ClustersDistributed TrainingToken Generation ThroughputRDMAInfiniBandMLPerfCheckpoint OptimizationDataset Streaming
Tech Stack
Tools & technologiesAnsibleAWSAzureGoogle Cloud PlatformGrafanaKubernetesLinuxNode.jsPrometheusPythonPyTorchTerraform
About the role
Key responsibilities & impact- Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads
- Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency
- Optimize inference clusters for token generation throughput, low latency, and high GPU utilization
- Build and support production AI infrastructure running hundreds to thousands of GPUs
- Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers
- Perform NCCL benchmarking, analysis, and tuning for collective communication performance
- Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS
- Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot
- Optimize GPU scheduling and resource allocation for training and inference environments
- Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases
- Identify performance regressions and troubleshoot distributed training issues at scale
- Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O
- Work closely with ML engineers to improve training scalability and inference efficiency
- Create automation to deploy, validate, benchmark, and monitor GPU clusters
- Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements
Requirements
What you’ll need- 7+ years designing or operating large-scale Linux infrastructure
- 5+ years supporting production GPU clusters for AI or HPC workloads
- Experience building multi-node GPU training environments from the ground up
- Deep expertise with distributed PyTorch training
- Extensive experience troubleshooting and optimizing NCCL communications
- Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications
- Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred
- Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism
- Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth
- Experience tuning CUDA, NCCL, UCX, and MPI
- Expert-level Linux systems administration skills
- Experience with Slurm
- Experience using Pyxis and Enroot for containerized GPU workloads
- Strong Python and Bash scripting skills
- Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet
- Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance
- Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling
- Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments
Benefits
Comp & perks- 🌐 Worldwide ❌ Jobs You've Hidden ⭐️ Saved Jobs ✅ Applied Jobs ✉️ Email Alerts 👤 Account STN Incorporated Website LinkedIn All Job Openings 11 - 50 employees Founded 2016 🏢 Enterprise 🔒 Cybersecurity 🔧 Hardware Enterprise
- Cybersecurity
- Hardware STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare. Cluster Engineer Job not on LinkedIn 🔥 0 minutes ago 🇺🇸 United States – Remote ⏰ Full Time 🟠 Senior 🔴 Lead 👷🏻♀️ Engineer Ansible AWS Azure Google Cloud Platform Grafana Kubernetes Linux Node.js Prometheus Python PyTorch Terraform Apply Now Find Hiring Managers Customize resume + cover letter Report problem ☆ Save ☑️ Mark as applied ❌ Hide 📋 Description
- Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads
- Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency
- Optimize inference clusters for token generation throughput, low latency, and high GPU utilization
- Build and support production AI infrastructure running hundreds to thousands of GPUs
- Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers
- Perform NCCL benchmarking, analysis, and tuning for collective communication performance
- Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS
- Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot
- Optimize GPU scheduling and resource allocation for training and inference environments
- Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases
- Identify performance regressions and troubleshoot distributed training issues at scale
- Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O
- Work closely with ML engineers to improve training scalability and inference efficiency
- Create automation to deploy, validate, benchmark, and monitor GPU clusters
- Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements 🎯 Requirements
- 7+ years designing or operating large-scale Linux infrastructure
- 5+ years supporting production GPU clusters for AI or HPC workloads
- Experience building multi-node GPU training environments from the ground up
- Deep expertise with distributed PyTorch training
- Extensive experience troubleshooting and optimizing NCCL communications
- Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications
- Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred
- Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism
- Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth
- Experience tuning CUDA, NCCL, UCX, and MPI
- Expert-level Linux systems administration skills
- Experience with Slurm
- Experience using Pyxis and Enroot for containerized GPU workloads
- Strong Python and Bash scripting skills
- Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet
- Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance
- Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling
- Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments Apply Now 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score Similar Jobs Deployed Engineer 🔥 30 minutes ago LangChain 11 - 50 🤖 Artificial Intelligence 🤝 B2B ☁️ SaaS Website LinkedIn All Job Openings Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform. 🇺🇸 United States – Remote 💵 $150k - $250k / year 💰 $25M Series A on 2024-02 ⏰ Full Time 🟡 Mid-level 🟠 Senior 👷🏻♀️ Engineer AWS Azure Google Cloud Platform JavaScript Kubernetes Python Forward Deployed Engineer 🔥 2 hours ago OnCorps AI 51 - 200 💸 Finance 💳 Fintech 🤖 Artificial Intelligence Website LinkedIn All Job Openings Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements. 🇺🇸 United States – Remote ⏰ Full Time 🟠 Senior 🔴 Lead 👷🏻♀️ Engineer AWS Docker Kubernetes NoSQL Pandas PySpark Python SQL TypeScript Fire Protection Engineer 🔥 2 hours ago Salas O'Brien 1001 - 5000 💼 Consulting 🏗️ Construction 🏥 Healthcare Website LinkedIn All Job Openings Fire protection engineer managing designs for hyperscale data centers, federal, mission-critical, and industrial projects at Salas O’Brien. Coordinating technical deliverables, calculations, Revit/BIM modeling, and junior-engineer guidance remotely. 🇺🇸 United States – Remote ⏰ Full Time 🟡 Mid-level 🟠 Senior 👷🏻♀️ Engineer Distinguished Engineer 🔥 2 hours ago Autodesk 10,000+ employees 🏗️ Construction 🏭 Manufacturing 💼 Consulting Website LinkedIn All Job Openings Distinguished Engineer leading Autodesk’s Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices. 🇺🇸 United States – Remote 💵 $196k - $352.1k / year ⏰ Full Time 🟠 Senior 🔴 Lead 👷🏻♀️ Engineer AWS Azure Cloud Distributed Systems Google Cloud Platform Kubernetes Microservices Strategic Engineer 1 – Night 🔥 3 hours ago NBCUniversal 10,000+ employees 📱 Media Website LinkedIn All Job Openings Strategic Engineer troubleshooting enterprise network and voice issues for Comcast Business, a connectivity and managed-solutions provider. Analyzing OSI-layer faults, configuring call flows, and escalating complex incidents. 🇺🇸 United States – Remote 💵 $26 / hour ⏰ Full Time 🟡 Mid-level 🟠 Senior 👷🏻♀️ Engineer 🦅 H1B Visa Sponsor View More Engineer Jobs 🌐 Worldwide Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or support@remoterocketship.com Search Search Jobs by country Search jobs by city Search jobs by job title Search entry-level jobs Search junior-level jobs Search senior-level jobs Search jobs by tech stack Search jobs by contract type Search remote internships Search remote part-time jobs Remote jobs Anywhere in the World Companies Hiring Anywhere in the World Companies Hiring Sales People Anywhere in the World Companies Hiring Software Engineers Anywhere in the World Resources Advice Tips for finding remote jobs Interview questions and answers Resume examples Cover letter examples Post a job Affiliates About us Is Remote Rocketship legit? Privacy policy Terms of service Job board SEO course Remote Job Search MasterClass AI Apply Copilot OpenClaw job finder Find jobs using your resume Jobs by Country Remote jobs anywhere in the world (Worldwide remote jobs) Remote jobs United States Remote jobs Australia Remote jobs Brazil Remote jobs Canada Remote jobs France Remote jobs Ireland Remote jobs Germany Remote jobs Netherlands Remote jobs Spain Remote jobs UK Popular Jobs Remote data analyst jobs Remote customer support jobs Remote executive assistant jobs Remote marketing jobs Remote product designer jobs Remote product manager jobs Remote project manager jobs Remote recruiter jobs Remote sales jobs Remote software engineer jobs Jobs by Type Remote full-time jobs Remote part-time jobs Remote contract jobs Remote internship jobs Remote entry-level jobs Remote jobs with no experience required Remote junior jobs (1-3 years of experience) Digital nomad jobs Remote jobs with no degree required Freelance remote jobs Temporary remote jobs Remote jobs hiring now Stay at home mom jobs