Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
OutcomesAI

Tech Lead – ASR, TTS, Speech LLM

OutcomesAI

Tech Lead specializing in ASR, TTS, and Speech LLM at OutcomesAI, a healthcare technology company. Leading technical development of speech models to enhance clinical capacity and efficiency.

Posted 7/29/2026full-timeSingapore • 🇸🇬 SingaporeSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates deep expertise in developing and deploying speech models, including ASR, TTS, and Speech LLM, with a strong focus on architecture, training strategies, and production readiness. Proven ability to mentor teams and guide technical roadmaps while ensuring high performance and accuracy in healthcare applications.

Highest-signal resume keywords
Speech Model DevelopmentASR/TTS Production ExperienceDeep Learning Frameworks (PyTorch, NeMo)Triton Inference Server DeploymentTelephony Robustness Expertise

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
ASRTTSSpeech LLMRNN-TCTC ArchitecturesLoRA/AdaptersTensorRT OptimizationSpeaker DiarizationEvaluation Metrics (WER, F1)Model Fine-Tuning
Soft Skills
MentorshipCollaborationCode Review Discipline
Tools & Technologies
Triton Inference ServerKubernetesGPU ScalingWFST GrammarsContext Injection
Certifications & Qualifications
M.S. / Ph.D. in Computer ScienceSpeech Processing
Industry Keywords
Healthcare ApplicationsSynthetic Data GenerationActive LearningInference OptimizationTelephony Noise

Tech Stack

Tools & technologies
KubernetesPyTorch

About the role

Key responsibilities & impact
  • Lead the end-to-end technical development of speech models (ASR, TTS, Speech-LLM) — from architecture, training strategy, and evaluation to production deployment.
  • Act as an individual contributor and mentor, guiding a small team working on model training, synthetic data generation, active learning, and inference optimization for healthcare applications.
  • Own the technical roadmap for STT/TTS/Speech LLM model training: from model selection → fine-tuning → deployment.
  • Evaluate and benchmark open-source models (Parakeet, Whisper, etc.) using internal test sets for WER, latency, and entity accuracy.
  • Design and review data pipelines for synthetic and real data generation (text selection, speaker selection, voice synthesis, noise/distortion augmentation).
  • Architect and optimize training recipes (LoRA/adapters, RNN-T, multi-objective CTC + MWER).
  • Lead integration with Triton Inference Server (TensorRT/FP16) and ensure K8s autoscaling for 1000+ concurrent streams.
  • Implement Language Model biasing APIs, WFST grammars, and context biasing for domain accuracy.
  • Guide evaluation cycles, drift monitoring, and model switcher/failover strategies.
  • Mentor engineers on data curation, fine-tuning, and model serving best practices.
  • Collaborate with backend/ML-ops for production readiness, observability, and health metrics.

Requirements

What you’ll need
  • M.S. / Ph.D. in Computer Science, Speech Processing, or related field.
  • 7–10 years of experience in applied ML, at least 3 in speech or multimodal AI.
  • Track record of shipping production ASR/TTS models or inference systems at scale.
  • Deep expertise in speech models (ASR, TTS, Speech LLM) and training frameworks (PyTorch, NeMo, ESPnet, Fairseq).
  • Proven experience with streaming RNN-T / CTC architectures, LoRA/adapters, and TensorRT optimization.
  • Telephony robustness: Codec augmentation (G.711 μ-law, Opus, packet loss/jitter), AGC/loudness norm, band-limit (300–3400 Hz), far-field/noise simulation.
  • Strong understanding of telephony noise, codecs, and real-world audio variability.
  • Experience in Speaker Diarization, turn detection model, smart voice activity detection Evaluation: WER/latency curves, Entity-F1 (names/DOB/meds), confidence metrics.
  • TTS: VITS/FastPitch/Glow-TTS/Grad-TTS/StyleTTS2, CosyVoice/NaturalSpeech-3 style transfer, BigVGAN/UnivNet vocoders, zero-shot cloning.
  • Speech LLM: Model development and integration with Voice agent pipeline.
  • Experience deploying models with Triton Inference Server, Kubernetes, and GPU scaling.
  • Hands-on with evaluation metrics (WER, F1 on entities, latency p50/p95).
  • Familiarity with LM biasing, WFST grammars, and context injection.
  • Strong mentorship and code-review discipline.

Benefits

Comp & perks
  • Health insurance
  • Professional development opportunities
  • Flexible working hours