FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Multimodal AI Model Optimization Research Engineer
TavusResearch Engineer focusing on optimizing AI models for multimodal interactions. Working with a collaborative team on groundbreaking human-AI interaction technologies.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in deep learning model optimization and compression techniques, including knowledge distillation and quantization, while effectively collaborating with researchers and engineers to develop deployable systems.
Highest-signal resume keywords
Deep Learning Using PyTorchModel Optimization and CompressionKnowledge DistillationQuantizationPython Coding Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model OptimizationKnowledge DistillationPruning/SparsificationQuantizationMixed PrecisionInference PerformanceGPU FundamentalsCloud EnvironmentsLarge ModelsResearch Engineering Practices
Soft Skills
Clear CommunicationCollaboration Skills
Industry Keywords
Efficient ArchitecturesLow-Rank AdaptersBenchmarkingLatencyCostQuality
Tech Stack
Tools & technologiesCloudPythonPyTorch
About the role
Key responsibilities & impact- Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization
- Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality
- Partner closely with researchers and engineers to turn new ideas into deployable systems
Requirements
What you’ll need- Strong experience in deep learning using PyTorch
- Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
- Understanding of efficient architectures such as low-rank adapters
- Strong understanding of inference performance and GPU/accelerator fundamentals
- Strong Python coding skills and reliable research engineering practices
- Experience working with large models and datasets in cloud environments
- Ability to read ML papers, reproduce results, and adapt ideas
- Clear communication and collaboration skills
Benefits
Comp & perks- flexible work schedules
- unlimited PTO
- competitive healthcare and gear stipends
- collaborative environment focused on learning and impact