FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Infrastructure Engineer
MirantisAI Infrastructure Engineer managing large-scale AI infrastructure systems for Mirantis, a Kubernetes-native AI company. Responsibilities include incident response, troubleshooting, and collaboration with a global team.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing and operating large-scale production systems, with a strong focus on Kubernetes, infrastructure automation, and troubleshooting complex technical issues. Proficient in scripting and familiar with NVIDIA GPU technologies to enhance operational efficiency.
Highest-signal resume keywords
Kubernetes ManagementInfrastructure AutomationTroubleshooting SkillsScripting ProficiencyHigh-Performance Computing
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesInfrastructure-as-CodeAnsibleTerraformPrometheusGrafanaELKPythonBashGo
Soft Skills
Analytical SkillsProblem-Solving SkillsEffective Communication
Tools & Technologies
Logging ToolsMonitoring ToolsProduction SystemsCloud EnvironmentsBare-Metal Environments
Industry Keywords
AI InfrastructureIncident ResponseRoot Cause AnalysisOperational DocumentationGlobal Collaboration
Tech Stack
Tools & technologiesAnsibleCloudGoGrafanaKubernetesPrometheusPythonTerraform
About the role
Key responsibilities & impact- Manage and operate production AI infrastructure environments.
- Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations.
- Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.
- Conduct root cause analysis and drive product and operational improvements.
- Contribute and improve operational documentation and knowledge base.
- Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.
Requirements
What you’ll need- Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).
- Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.
- Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
- Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
- Effective verbal and written communication skills in English.
- Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.
- Willingness to occasionally work weekends and holidays.
- Nice to have: Previous experience building, scaling, and running High-Performance Computing (HPC) environments.
- Hands-on experiences with managing large scale Kubernetes platforms in production.
- A good understanding of NVIDIA GPU technologies and the associated software stack.
- Proficiency in scripting languages (e.g., Python, Bash, Go).
Benefits
Comp & perks- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.