Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
FluidStack

Principal Operations Engineer, Reliability

FluidStack

Principal Operations Engineer responsible for fleet reliability at Fluidstack's AI compute infrastructure. Defining availability targets and implementing corrective actions to improve operations.

Posted 7/28/2026full-timeRemote • 🇺🇸 United StatesLead💰 $220,000 - $260,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in reliability engineering for critical infrastructure, focusing on availability targets and root cause analysis. Proficient in utilizing failure data to inform maintenance strategies and drive corrective actions across teams.

Highest-signal resume keywords
Reliability EngineeringRoot Cause AnalysisFailure Data AnalysisCorrective Action ImplementationCMMS Analytics

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Availability MeasurementReliability-Centered MaintenanceCondition-Based MaintenanceWeibull AnalysisPareto AnalysisFMEAData Pipeline DevelopmentIncident History AnalysisCritical Infrastructure ManagementFailure Data Utilization
Soft Skills
Cross-Team CollaborationAnalytical ThinkingProblem Solving
Tools & Technologies
Data Center ReliabilityPower Generation ReliabilityLiquid Cooling Systems
Certifications & Qualifications
CRE CertificationCMRP Certification

About the role

Key responsibilities & impact
  • Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap.
  • Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site.
  • Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities.
  • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.

Requirements

What you’ll need
  • You've owned reliability for critical infrastructure and moved the availability number, not just reported it.
  • You've led root cause analyses that found the real cause, not the convenient one.
  • You work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use, not terms you know.
  • You get corrective actions closed across teams you don't manage.
  • Bonus: Data center or power generation reliability. Liquid cooling systems. CMMS analytics. CRE or CMRP certification.

Benefits

Comp & perks
  • base salary
  • equity for all full time roles
  • benefits
  • commissions plans if applicable