Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CloudLinux

Lead Site Reliability Engineer – Imunify Reliability Platform

CloudLinux

Lead SRE defining SLIs, SLOs, alerting, and escalation for CloudLinux’s Imunify360 Linux security platform. Building telemetry and reliability practices across roughly 70 components.

Posted 8/22/2026full-timeRemote • 🇵🇱 PolandSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in defining and implementing SLI frameworks, telemetry collection pipelines, and alerting systems while ensuring effective communication and collaboration across teams. Proficient in Python, Go, and Rust for building and maintaining production-scale systems.

Highest-signal resume keywords
Production Engineering ExperienceSLI Framework DefinitionStrong Python SkillsTelemetry Collection Pipeline DesignPrometheus/OpenMetrics Experience

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SLI DefinitionSLO FrameworkTelemetry CollectionDistributed Systems DebuggingConfiguration ManagementCI/CD PracticesTime-Series TelemetryAlerting Systems DesignError Budget ManagementData Analysis
Soft Skills
Strong Written CommunicationTeam CollaborationIncident Command Practice
Tools & Technologies
PrometheusGrafanaAlertmanagerClickHouseAnsibleGitLab CIJenkins
Industry Keywords
Production EngineeringSite Reliability EngineeringTelemetryService Level IndicatorsService Level ObjectivesAlertingIncident Management

Tech Stack

Tools & technologies
AnsibleDistributed SystemsGoGrafanaJenkinsPrometheusPythonRust

About the role

Key responsibilities & impact
  • Define what “working” means for approximately 70 components
  • Run SLI definition with squad leads and senior engineers and facilitate ownership sign-off
  • Build service, fleet, control-efficacy, delivery, and pipeline SLI taxonomies
  • Attach an SLO, error budget, and owning squad to each SLI, with appropriate tiering
  • Design and build a push-based, sampled, privacy-constrained telemetry collection pipeline with a defended cardinality budget
  • Extend agent-side and service-side instrumentation in Python, Go, and Rust
  • Consolidate dashboards, ad-hoc queries, and reporting paths into a defensible instrument set
  • Build symptom-based, SLO-anchored alerting with multi-window burn-rate semantics
  • Establish page, ticket, and dashboard alert tiers and define paging criteria
  • Ensure every alert has an owner, runbook, and documented failure mode
  • Maintain alert hygiene through quarterly reviews, deletion, and actionable-rate tracking
  • Maintain a machine-readable component-to-squad ownership map wired into alert routing
  • Design severity matrices, acknowledgement SLAs, follow-the-sun rotas across UTC−5 to UTC+8, and handoff protocols
  • Practice incident command and blameless postmortems within 24 hours
  • Design escalation so squads carry their own pagers while you operate the platform and coach teams
  • Deliver first-year outcomes including component inventory, pilot instrumentation, production collection pipeline, tier-1 escalation, full component SLI coverage, alert metrics, squad on-call, and mean time to detect silent control degradation under 24 hours

Requirements

What you’ll need
  • Substantial production-engineering or SRE experience, including at least one environment where you defined the SLO framework rather than inherited it
  • Ability to walk through personally written SLIs and explain how they were negotiated with resistant teams
  • Strong Python
  • Comfortable reading and modifying Go or Rust
  • Practical experience with time-series and event telemetry at scale
  • Prometheus/OpenMetrics, Grafana, and an Alertmanager-class routing layer
  • Experience with ClickHouse or an equivalent columnar store for high-cardinality fleet data
  • Distributed systems debugging on bare metal and long-lived hosts
  • Configuration management and CI at production scale, including Ansible, GitLab CI, Jenkins, or close equivalents
  • Ability to design measurement for machines that cannot be owned or scraped, including push telemetry, sampling, clock skew, partial reporting, and customer-server privacy constraints
  • Strong written communication suitable for async work

Benefits

Comp & perks
  • Professional development opportunities
  • Interesting and challenging projects
  • Mentor and knowledge-exchange programs
  • Fully remote work with flexible working hours
  • Work from any location worldwide
  • 24 days of paid vacation per year
  • 10 days of national holidays
  • Unlimited sick leave
  • Compensation for private medical insurance
  • Co-working reimbursement
  • Gym/sports reimbursement
  • Opportunity to receive a reward for the most innovative idea that the company can patent