Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cantina

Machine Learning Engineer – Voice Conversion

Cantina

Machine Learning Engineer building state-of-the-art speech and voice-conversion systems for Cantina’s AI social platform. Owning models, data, evaluation, and production inference pipelines.

Posted 8/5/2026full-timeRemote • 🇺🇸 United StatesMid-LevelSenior💰 $200,000 - $220,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting and implementing large-scale speech models, with a strong focus on training methodologies, data management, and system optimization. Proficient in collaborating across teams to enhance speech technology and ensure responsible AI practices.

Highest-signal resume keywords
Large-Scale Audio ModelsDiffusion TransformersMulti-Node Distributed TrainingPyTorch ProficiencyVoice Cloning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Large-Scale Audio ModelsDiffusion TransformersFlow-Matching TransformersAudio VAEsNeural Audio CodecsMulti-Node Distributed TrainingProduction-Quality CodePerformance ProfilingCUDAAdversarial Training
Tools & Technologies
FSDPDeepSpeedPyTorchCUDATriton
Industry Keywords
Speech TechnologyMachine LearningGenerative ModelsData CurationModel Evaluation

Tech Stack

Tools & technologies
Node.jsPyTorch

About the role

Key responsibilities & impact
  • Architect, implement, pre-train, fine-tune, and post-train/alignment large-scale speech models
  • Design, run, and analyze scientific experiments to advance understanding of models
  • Develop and improve development tooling to enhance team productivity
  • Contribute across the stack, from low-level optimizations to high-level model design
  • Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies
  • Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR-based metrics, robustness and bias checks, and red-team studies
  • Harden the training-to-evaluation-to-inference pipeline; profile latency, memory, and cost; meet production SLAs with monitoring and rollback
  • Contribute to safety and consent guardrails and misuse/abuse mitigation for responsible speech technology
  • Partner with research, data, infrastructure, and product teams to ship measurable improvements to speech systems

Requirements

What you’ll need
  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data)
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning mechanisms, and distillation
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training
  • Strong experience with multi-node, multi-GPU distributed training using FSDP, DeepSpeed, or equivalent
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong proficiency with PyTorch and performance work, including profiling and CUDA/Triton/C++ as needed
  • Experience writing reliable production-quality code
  • Experience shipping large-scale speech/audio or multimodal generative models to production
  • Background working with large-scale ML data and triangulating quality using subjective and objective signals
  • Experience with voice cloning, speech control/steerability, or expressive speech generation
  • Notable publications and/or open-source contributions in speech, audio, or machine learning

Benefits

Comp & perks
  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!