FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Compute Deployment Engineer
FluidStackCompute Deployment Engineer maintaining server and GPU fleets at scale, ensuring efficient production and automated workflows. Collaborating across teams during system turn-up and incident responses.
Posted 7/21/2026full-timeNew York City • New York • 🇺🇸 United StatesMid-LevelSenior💰 $197,000 - $227,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in server and GPU fleet turn-up, including automation of hardware workflows and effective triage of hardware failures. Proficient in Linux and out-of-band management tools, with hands-on experience in data center operations and remote support.
Highest-signal resume keywords
Server Fleet Turn-UpLinux ProficiencyHardware Workflow AutomationKubernetes-Based ProvisioningTriage Hardware Failures
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
BMCIPMIRedfishPythonGoData Hall OperationsBurn-In TestingStress Harness DesignComponent SwappingFirmware Updates
Tools & Technologies
KubernetesDCIMInventory ToolingCustom Accelerator Platforms
Industry Keywords
GPUData Center OperationsRemote HandsRMAVendor Escalation
Tech Stack
Tools & technologiesGoKubernetesLinuxNode.jsPython
About the role
Key responsibilities & impact- Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
- Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
- Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
- Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
- Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
- Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Requirements
What you’ll need- You've brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
- You work deep in Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.
- You've automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
- You've worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.
- You triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
- You travel for turn-up windows when a new data hall comes online.
- Bonus: Kubernetes-based bare-metal provisioning. Accelerator platform bringup (NVIDIA, AMD, or custom). Burn-in and stress harness design. DCIM and inventory tooling.
Benefits
Comp & perks- Offers Equity 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score