FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Applied Scientist – Multimodal
FlawlessApplied Scientist developing scalable systems for audio/video dataset curation and lip sync model training. Collaborating with teams to enhance model performance and quality metrics in entertainment AI.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing scalable audio/video dataset curation pipelines and lip sync model training workflows, with a strong foundation in Python and deep learning frameworks like PyTorch. Proficient in audio processing, 3D computer vision, and speech synthesis, with a focus on delivering value through effective collaboration and communication.
Highest-signal resume keywords
Python ProgrammingDeep Learning Frameworks (PyTorch)Audio Processing3D Computer VisionSpeech Synthesis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model DesignData ProcessingStatistical Methods for Signal ProcessingAudio-Visual LearningMultimodal FusionSpeech ProcessingDialog/Speaker DetectionSpeaker SeparationLip Sync Model TrainingQuantitative and Qualitative Metrics
Soft Skills
Outstanding Communication SkillsCollaboration
Tools & Technologies
OpenCV
Industry Keywords
Audio/VisualMultimodal Related FieldsVFXResearch Initiatives
Tech Stack
Tools & technologiesPythonPyTorch
About the role
Key responsibilities & impact- Develop repeatable, scalable audio/video dataset curation pipelines and lip sync model training workflows across multiple datasets
- Train, fine-tune, and manage audio/video and lip sync model variants as model dependencies, data, and architectures evolve
- Incorporate new datasets and model updates as they become available
- Design, automate, and maintain audio/video datasets and lip sync metric testing pipelines
- Generate new quantitative and qualitative metrics to evaluate audio/video and lip sync quality
- Produce comparisons, visualizations, and analyses to inform research and product decisions
- Partner closely with audio/video and lip sync researchers to support ongoing and future research initiatives
- Validate audio/video and lip sync quality to improve out-of-the-box approval rates and reduce downstream cost and iteration time
- Collaborate with Science, Engineering, and Product teams to align research outputs with company goals
Requirements
What you’ll need- MSc OR PhD + Industry experience working in the domains of Audio processing, 3D Computer Vision, Speech Synthesis, Computer Graphics, or other multimodal related fields such as text/audio, or audio/visual.
- Proficiency in Python, with a strong foundation in computer science and problem-solving.
- Expertise with deep learning frameworks (PyTorch) and vision tools (OpenCV).
- A strong product mindset — motivated by building systems that deliver tangible value to users, not just technical novelty.
- Comfortable working at both the algorithmic and implementation levels, from model design and optimisation to large-scale data processing and integration in production systems.
- High degree of proficiency in math and statistical methods for signal processing
- Experience with audio-visual learning, multimodal fusion, and/or audio-driven face animation
- Experience with speech processing and detection, such as dialog/speaker detection, speaker separation, and speech synthesis with deep neural networks
- Outstanding communication skills for collaboration with scientists, research/ML engineers, and VFX artists.
Benefits
Comp & perks- Autonomy
- A hybrid working environment
- Competitive Salary
- All permanent employees receive generous stock options