Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Yotta Labs

Research Engineer Intern – AI Systems

Yotta Labs

Research Engineer Intern focused on AI Systems at Yotta Labs. Work on optimizing GPU kernels and LLM infrastructure in a remote setting.

Posted 8/2/2026internshipRemote • 🇺🇸 United StatesEntry LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in implementing and optimizing compute kernels for AI frameworks, with strong programming skills in Python and experience in GPU architecture. Proficient in building custom operators and profiling performance in collaborative environments.

Highest-signal resume keywords
CUDA ProgrammingPyTorch FrameworkGPU Architecture FundamentalsPerformance Profiling ToolsCustom Operator Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Python ProgrammingC++ ProgrammingCUDATritonROCm/HIPNeuron SDKAI FrameworksModel ArchitecturesKernel OptimizationMemory Optimization
Soft Skills
Problem-SolvingCollaborationIndependence
Tools & Technologies
NsightROCm ProfilerNeuron Profiler
Industry Keywords
Compute KernelsAttentionGEMMMoEQuantizationInference PerformanceBenchmarkingOpen-Source AI Infrastructure

Tech Stack

Tools & technologies
AWSPythonPyTorch

About the role

Key responsibilities & impact
  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.
  • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
  • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations.
  • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.
  • Ship code upstream to open-source AI infrastructure projects, with tests and documentation.

Requirements

What you’ll need
  • Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field
  • Solid programming skills in Python and familiarity with C++
  • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects
  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count
  • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler)
  • Strong problem-solving skills and the ability to work independently in a collaborative, remote environment.

Benefits

Comp & perks
  • Competitive internship compensation
  • Flexible remote work environment
  • Direct mentorship from engineers from leading institutions and tech companies
  • Fast path to a full-time return offer for top performers