Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Runware

Senior Site Reliability Engineer

Runware

Site Reliability Engineer ensuring reliability and performance of Runware's AI platforms. Collaborating across software, infrastructure, and operations to enhance observability and reduce incidents.

Posted 7/20/2026full-timeRemote • 🇬🇧 United KingdomSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in operating and troubleshooting production systems at scale, with a strong focus on reliability practices, observability, and incident management. Proficient in automation and improving system resilience through effective engineering solutions.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Distributed Systems DebuggingKubernetes and Container ManagementObservability Systems DesignIncident Management and RCA

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production Systems OperationAutomation and RemediationCapacity PlanningError BudgetsPython ProgrammingGo ProgrammingPHP ProgrammingDatabase Management (MySQL, Redis, ClickHouse)Distributed Messaging Systems (RabbitMQ)Infrastructure as Code (IaC)
Soft Skills
Ownership of Production ProblemsCollaboration with Engineering Teams
Tools & Technologies
KubernetesContainersCDN PlatformsLoad BalancingHybrid Infrastructure
Industry Keywords
SLIsSLOsIncident ManagementOperational Toil ReductionHigh-Throughput APIsLow-Latency APIsGPU EnvironmentsAI and ML Workloads

Tech Stack

Tools & technologies
Distributed SystemsGoKubernetesMySQLPHPPythonRabbitMQRedis

About the role

Key responsibilities & impact
  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements

What you’ll need
  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
  • Bonus
  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

Benefits

Comp & perks
  • Generous paid time off – vacation, sick days, public holidays
  • Meaningful stock options – share in the upside you create
  • Remote-first setup – work from home anywhere we can employ you
  • Flexible hours – own your schedule outside core collaboration blocks
  • Family leave – paid maternity, paternity, and caregiver time
  • Company retreats – twice-yearly gatherings in inspiring locations