Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Tubi

Software Engineer, ML Infra, Distributed Systems – Staff, Principal

Tubi

Software Engineer developing scalable and low-latency ML infrastructure for Tubi's machine learning platform. Collaborating closely with ML and Product teams to enhance user experiences.

Posted 7/27/2026full-timeToronto • 🇨🇦 CanadaLead💰 CA$164,600 - CA$268,900 per yearWebsite

Tech Stack

Tools & technologies
AWSCassandraCloudDistributed SystemsDockerGoJavaKafkaKubernetesMicroservicesNoSQLPostgresPythonRedisScalaSQL

About the role

Key responsibilities & impact
  • Design and build scalable, high throughput, and low latency distributed systems using Scala
  • Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration
  • Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art.
  • Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary
  • Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc.
  • Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi.

Requirements

What you’ll need
  • Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM based language is a plus.
  • Strong experience with AWS or an equivalent cloud platform
  • Experience building online microservices at scale with low latency serving
  • Experience with both SQL (e.g. Postgres) and NoSQL databases (e.g. Cassandra), message brokers (e.g. Kafka), and caches (e.g. Redis)
  • Experience with containerization technologies, such as Docker or Kubernetes
  • Led the response and resolution efforts for multiple major, large-scale incidents.

Benefits

Comp & perks
  • Medical/dental/vision insurance
  • Vacation/paid time off
  • Annual discretionary bonus
  • Long-term incentive plan
  • Monthly wellness reimbursement
  • Flexible Time Off Policy
  • Generous Parental Leave Program