FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Infrastructure Engineer
42dotAI Infrastructure Engineer managing high-performance AI infrastructure orchestrating GPUs at 42dot. Contributing to scaling, monitoring, and operational optimization of a world-class computing environment.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing and maintaining large-scale GPU clusters using Kubernetes and Slurm, with a strong focus on automation through Python or Shell scripting. Proficient in Linux operating systems and containerization technologies, ensuring optimal resource utilization and effective communication across teams.
Highest-signal resume keywords
Kubernetes ManagementGPU Cluster OperationsPython ScriptingLinux Operating SystemsDocker Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesSlurmPythonShell ScriptingLinuxDockerNetworking FundamentalsTCP/IPHTTP(S)Problem Solving
Soft Skills
Communication SkillsCollaboration
Industry Keywords
GPU Resource ManagementHigh AvailabilityDistributed Learning EnvironmentAutomation ToolsSystem Administration
Tech Stack
Tools & technologiesDockerKubernetesLinuxPythonTCP/IP
About the role
Key responsibilities & impact- Kubernetes 및 Slurm을 활용하여 여러 데이터 센터에 분산된 수천 개 규모의 대규모 GPU 클러스터 운영 및 유지 보수
- GPU 하드웨어 및 소프트웨어 스택 전반의 장애를 모니터링하고 진단하여 고가용성 유지 및 신속한 장애 복구 수행
- Python 또는 Shell을 활용한 자동화 도구 및 스크립트를 개발하여 반복적인 인프라 관리 업무를 효율화
- GPU 리소스 쿼터(Quota) 관리 및 ML 개발자를 위한 기술 지원을 통해 컴퓨팅 자원의 최적 활용 보장
- 대규모 자율주행 모델 학습을 위한 분산 학습 환경의 아키텍처 설계 및 성능 튜닝 참여.
Requirements
What you’ll need- Linux 운영체제에 대한 깊은 이해 (커널 동작, 프로세스 관리, 시스템 보안 등)
- Docker 및 Kubernetes 등 컨테이너 기반 기술 및 오케스트레이션 실무 경험
- TCP/IP, HTTP(S) 등 네트워크 기본 원리에 대한 이해 및 기초적인 네트워크 트러블슈팅 능력
- Python 또는 Shell을 활용하여 유지보수가 용이한 자동화/시스템 관리 스크립트 작성 역량
- 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력
- 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량.
- Strong proficiency in Linux operating systems, including a solid understanding of kernel operations, process management, and system security.
- Practical experience with containerization technologies (Docker) and orchestration (Kubernetes), including building, managing, and troubleshooting containerized environments.
- Solid understanding of networking fundamentals, including TCP/IP and HTTP(S), with the ability to perform basic network troubleshooting.
- Ability to write clean and maintainable scripts in Python or Shell for automation and system administration.
- Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems.
- Strong communication skills to effectively collaborate with cross-functional teams and external partners.
Benefits
Comp & perks- 전형 절차는 일정 및 진행 상황에 따라 일부 변경될 수 있으며, 각 전형 결과는 등록하신 이메일로 개별 안내드립니다.
- 지원서 제출 시 주민등록번호, 가족관계, 혼인 여부, 연봉, 사진, 신체조건, 출신 지역 등 채용절차법상 요구 금지된 정보는 제외 부탁드립니다.
- 지원서 접수 중 오류가 발생하거나 기타 문의 사항이 있을 경우, recruit@42dot.ai로 문의해 주시기 바랍니다.
- 국가보훈대상자 및 취업보호 대상자는 관계법령에 따라 우대합니다.
- 장애인 고용 촉진 및 직업재활법에 따라 장애인 등록증 소지자를 우대합니다.