Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
42dot

LLM Engineer – Optimization

42dot

LLM Engineer optimizing inference performance for AI systems across various environments. Focusing on model optimization, performance analysis, and inference engine development.

Posted 7/6/2026full-timePangyo • 🇰🇷 South KoreaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Expertise in optimizing LLM Inference Performance through techniques such as Quantization, Model Compression, and Compiler Optimization, with proficiency in deep learning frameworks like PyTorch and TensorRT. Strong analytical and problem-solving skills, along with collaboration capabilities in high-performance computing environments.

Highest-signal resume keywords
LLM Inference Engine DevelopmentGPU Architecture UnderstandingCUDA ProgrammingQuantization TechniquesDeep Learning Frameworks Experience

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Inference OptimizationModel CompressionCompiler OptimizationQuantizationPerformance AnalysisPython ProgrammingC/C++ ProgrammingParallel ComputingMachine Learning InfrastructureBenchmarking
Soft Skills
Problem-SolvingCollaboration
Tools & Technologies
PyTorchONNXTensorRTVLLMSGLangLlama.cppONNX RuntimeMLX
Industry Keywords
Large Language ModelInference FrameworkGPUEmbedded SystemsEdge Device

Tech Stack

Tools & technologies
C++PythonPyTorch

About the role

Key responsibilities & impact
  • 대규모 언어 모델(LLM)의 추론 성능(Latency, Throughput, Memory Efficiency)을 최적화합니다.
  • 다양한 모델 구조 및 추론 환경에 맞는 최적화 기법을 연구하고 적용합니다.
  • Long Context, Multi-turn Conversation 등 실제 서비스 환경에서의 성능을 개선합니다.
  • GPU 및 Accelerator 기반 LLM Inference Engine을 개발하고 최적화합니다.
  • vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX 등 최신 Inference Framework를 활용하거나 개선합니다.
  • Quantization(MXFP8, NVFP4, AWQ, GPTQ 등) Pruning, Distillation 등 모델 경량화 기법을 연구하고 적용합니다.
  • Mobile, Embedded, Edge Device 환경에서 LLM을 효율적으로 실행하기 위한 최적화 기술을 개발합니다.
  • 다양한 하드웨어 및 Inference Backend의 성능을 분석하고 Benchmark를 수행합니다.

Requirements

What you’ll need
  • ILLM, Machine Learning Infrastructure 또는 Inference Optimization 관련 경력 3년 이상
  • LLM Inference Engine 또는 AI Runtime 개발 경험
  • GPU Architecture, CUDA Programming 또는 병렬 컴퓨팅에 대한 이해
  • Quantization, Model Compression, Compiler Optimization 등 모델 최적화 기술에 대한 이해
  • PyTorch, ONNX, TensorRT 등 딥러닝 프레임워크 활용 경험
  • Python 및 C/C++ 중 하나 이상의 언어에 능숙하며, 소프트웨어 엔지니어링 역량을 보유하신 분
  • 성능 분석 및 문제 해결 능력과 협업 역량을 갖추신 분.

Benefits

Comp & perks
  • 국가보훈대상자 및 취업보호 대상자는 관계법령에 따라 우대합니다.
  • 장애인 고용 촉진 및 직업재활법에 따라 장애인 등록증 소지자를 우대합니다.
  • 42dot은 의뢰하지 않은 서치펌의 이력서를 받지 않으며, 요청하지 않은 이력서에 대해 수수료를 지불하지 않습니다.
  • 지원서 내용 중 허위 사실이 발견될 경우, 입사가 취소될 수 있습니다.
  • 인터뷰 프로세스 종료 후 지원자의 동의하에 평판조회가 진행될 수 있습니다.
  • 3개월의 수습기간이 적용될 수 있습니다.