Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

AI Computing Software Development Engineer, TensorRT-LLM

NVIDIA

Software engineer developing NVIDIA TensorRT-LLM inference software for GPU-accelerated AI. Optimizing LLM performance, kernels, runtimes, and deployment across platforms.

Posted 8/10/2026full-timeTaipei • 🇹🇼 TaiwanMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and optimizing Large Language Model (LLM) inference software, with strong proficiency in Python and C/C++. Capable of collaborating across teams to enhance deep learning inference performance and architecture.

Highest-signal resume keywords
Python ProgrammingC/C++ ProgrammingDeep Learning FrameworksLarge Language Model InferenceGPU Programming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software DevelopmentSoftware DesignPerformance ModelingCode OptimizationDebuggingTest DesignInference AlgorithmsArchitectural KnowledgeRuntime FeaturesKernels Implementation
Soft Skills
Excellent Communication SkillsProactive Work Ethic
Tools & Technologies
TensorRT-LLMPyTorchHuggingFaceCUDAOpenCL
Industry Keywords
Artificial IntelligenceDeep LearningHigh-Performance ComputingLLM ArchitecturesInference Frameworks

Tech Stack

Tools & technologies
PythonPyTorch

About the role

Key responsibilities & impact
  • Craft and develop robust inference software scalable across multiple platforms for functionality and performance
  • Analyze and optimize Large Language Model (LLM) inference performance
  • Follow academic and industrial developments in artificial intelligence and feature-update TensorRT-LLM
  • Implement kernels and runtime features supporting new LLM models and inference algorithms
  • Provide feedback on architecture and hardware design and development
  • Collaborate across software, research, and product teams to guide deep learning inference direction

Requirements

What you’ll need
  • Master or higher degree in Computer Engineering, Computer Science, Applied Mathematics, or related computing-focused degree, or equivalent experience
  • 3+ years of relevant software development experience
  • Excellent Python programming, software design, and software engineering skills
  • Awareness of current LLM architectures and LLM inference techniques
  • Experience with deep learning frameworks such as PyTorch and HuggingFace
  • Ability to work proactively and without supervision
  • Excellent written and oral communication skills in English
  • Prior experience with an LLM inference framework or a deep-learning compiler is preferred
  • Understanding of operations inside LLM inference frameworks or the end-to-end LLM inference workflow is preferred
  • Experience in performance modeling, profiling, debugging, and code optimization of DL/HPC/high-performance applications is preferred
  • Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design, are preferred
  • Architectural knowledge of CPU and GPU is preferred
  • GPU programming experience with CUDA or OpenCL is preferred

Benefits

Comp & perks
  • Competitive salaries
  • Generous benefits package
  • Equal opportunity employer
  • Diversity valued