Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Mirantis

AI Infrastructure Engineer

Mirantis

AI Infrastructure Engineer managing large-scale AI infrastructure systems for Mirantis, a Kubernetes-native AI company. Responsibilities include incident response, troubleshooting, and collaboration with a global team.

Posted 7/29/2026full-timeRemote • 🇯🇵 JapanMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing and operating large-scale production systems, with a strong focus on Kubernetes, infrastructure automation, and troubleshooting complex technical issues. Proficient in scripting and familiar with NVIDIA GPU technologies to enhance operational efficiency.

Highest-signal resume keywords
Kubernetes ManagementInfrastructure AutomationTroubleshooting SkillsScripting ProficiencyHigh-Performance Computing

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesInfrastructure-as-CodeAnsibleTerraformPrometheusGrafanaELKPythonBashGo
Soft Skills
Analytical SkillsProblem-Solving SkillsEffective Communication
Tools & Technologies
Logging ToolsMonitoring ToolsProduction SystemsCloud EnvironmentsBare-Metal Environments
Industry Keywords
AI InfrastructureIncident ResponseRoot Cause AnalysisOperational DocumentationGlobal Collaboration

Tech Stack

Tools & technologies
AnsibleCloudGoGrafanaKubernetesPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Manage and operate production AI infrastructure environments.
  • Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations.
  • Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.
  • Conduct root cause analysis and drive product and operational improvements.
  • Contribute and improve operational documentation and knowledge base.
  • Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.

Requirements

What you’ll need
  • Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).
  • Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.
  • Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
  • Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
  • Effective verbal and written communication skills in English.
  • Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.
  • Willingness to occasionally work weekends and holidays.
  • Nice to have: Previous experience building, scaling, and running High-Performance Computing (HPC) environments.
  • Hands-on experiences with managing large scale Kubernetes platforms in production.
  • A good understanding of NVIDIA GPU technologies and the associated software stack.
  • Proficiency in scripting languages (e.g., Python, Bash, Go).

Benefits

Comp & perks
  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.