Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Opedia Technologies

Data Center Operations and Maintenance Engineer

Opedia Technologies

Operations & Maintenance Engineer supporting next-generation AI infrastructure for a cloud platform. Collaborating with cross-functional teams to ensure reliability and customer support.

Posted 7/21/2026full-timeBellevue • Washington • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in monitoring and maintaining data center and GPU infrastructure, with a strong focus on incident response, troubleshooting, and operational maturity. Collaborates effectively with cross-functional teams to enhance monitoring and alerting systems, ensuring high reliability and performance.

Highest-signal resume keywords
Data Center OperationsIncident ResponseTroubleshooting SkillsGPU Infrastructure SupportOperational Maturity

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
MonitoringAlertingTroubleshootingCapacity ManagementPerformance Monitoring
Soft Skills
CollaborationProblem-SolvingAdaptability
Tools & Technologies
Operational ToolingRunbooksOn-Call Practices
Industry Keywords
Site ReliabilityInfrastructure OperationsLarge-Scale Compute Environments

About the role

Key responsibilities & impact
  • Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues.
  • Lead or support incident response for production issues, driving toward fast, effective resolution.
  • Collaborate with the engineering teams to make sure monitoring, alerting, and operational tooling are first class and make issues are caught before they impact customers.
  • Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence.
  • Contribute to runbooks, on-call practices, and operational maturity as the platform scales.

Requirements

What you’ll need
  • Experience in a data center operations, site reliability, or infrastructure operations role, ideally supporting GPU or large-scale compute environments.
  • Strong incident response and troubleshooting skills across hardware, networking, and systems layers.
  • Comfortable working in an early-stage environment where processes and tooling are still being established.
  • Experience in multiple technical environments is a plus.

Benefits

Comp & perks
  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance.
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.