Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
FluidStack

Compute Deployment Engineer

FluidStack

Compute Deployment Engineer maintaining server and GPU fleets at scale, ensuring efficient production and automated workflows. Collaborating across teams during system turn-up and incident responses.

Posted 7/21/2026full-timeNew York City • New York • 🇺🇸 United StatesMid-LevelSenior💰 $197,000 - $227,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in server and GPU fleet turn-up, including automation of hardware workflows and effective triage of hardware failures. Proficient in Linux and out-of-band management tools, with hands-on experience in data center operations and remote support.

Highest-signal resume keywords
Server Fleet Turn-UpLinux ProficiencyHardware Workflow AutomationKubernetes-Based ProvisioningTriage Hardware Failures

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
BMCIPMIRedfishPythonGoData Hall OperationsBurn-In TestingStress Harness DesignComponent SwappingFirmware Updates
Tools & Technologies
KubernetesDCIMInventory ToolingCustom Accelerator Platforms
Industry Keywords
GPUData Center OperationsRemote HandsRMAVendor Escalation

Tech Stack

Tools & technologies
GoKubernetesLinuxNode.jsPython

About the role

Key responsibilities & impact
  • Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
  • Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
  • Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
  • Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
  • Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
  • Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.

Requirements

What you’ll need
  • You've brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
  • You work deep in Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.
  • You've automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
  • You've worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.
  • You triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
  • You travel for turn-up windows when a new data hall comes online.
  • Bonus: Kubernetes-based bare-metal provisioning. Accelerator platform bringup (NVIDIA, AMD, or custom). Burn-in and stress harness design. DCIM and inventory tooling.

Benefits

Comp & perks
  • Offers Equity 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score