FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Observability Platform Engineer
MirantisObservability Platform Engineer building metrics, logging, tracing, and alerting platforms for Mirantis’s Kubernetes-native AI infrastructure. Improving incident detection and resolution across globally distributed compute environments.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and building observability platforms for large-scale production environments, with a strong focus on metrics, logging, and distributed tracing. Proficient in implementing SLO/SLI frameworks and integrating observability tooling with incident management workflows.
Highest-signal resume keywords
Observability Platform DesignMetrics, Logging, Distributed TracingSLO/SLI Framework ImplementationHigh-Volume Telemetry PipelinesKubernetes and Cloud-Native Infrastructure
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GoPythonRustPrometheusGrafanaOpenTelemetryLokiElasticsearchJaegerEBPF
Soft Skills
Strong Communication SkillsAbility to Work in Fast-Moving Environment
Tools & Technologies
ThanosCortexMimirAIOpsAutomated Triage Systems
Industry Keywords
High-Performance ComputeManaged ServicesOpen-Source ContributionsIncident Management WorkflowsRoot-Cause Analysis
Tech Stack
Tools & technologiesCloudElasticSearchGoGrafanaKubernetesPrometheusPythonRust
About the role
Key responsibilities & impact- Design, build, and operate metrics, logging, distributed tracing, and alerting platform components
- Build high-volume, high-cardinality telemetry pipelines for large infrastructure fleets
- Define and implement SLO/SLI frameworks and alerting strategies
- Partner with service delivery and operations teams to build incident-focused observability
- Integrate observability tooling with incident management workflows, root-cause analysis, and post-incident reviews
- Improve detection speed and reduce MTTD and MTTR
- Contribute to AI-assisted operations tooling, including automated triage, anomaly detection, and engineer-assist tools
- Own the reliability, scalability, and security of the observability stack
- Document architecture, runbooks, and operational practices
Requirements
What you’ll need- Proven experience designing and building observability platforms for large-scale, production infrastructure environments
- Strong hands-on experience with metrics, logging, and distributed tracing tooling, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, and Jaeger/Tempo, or equivalents
- Experience with high-volume telemetry pipelines and tradeoffs involving cardinality, retention, cost, and query latency
- Strong software engineering skills in at least one relevant language, such as Go, Python, or Rust
- Experience with Kubernetes and cloud-native infrastructure
- Solid understanding of SLO/SLI/error-budget practices and low-noise alerting design
- Comfortable working in a fast-moving environment
- Strong communication skills and ability to work directly with operations/service delivery teams
- Preferred: observability for GPU/HPC or specialized high-performance compute environments
- Preferred: eBPF-based observability tooling
- Preferred: AIOps/ML-based anomaly detection or automated triage systems
- Preferred: managed services or MSP context
- Preferred: contributions to open-source observability projects
Benefits
Comp & perks- Professional development and training
- Attend conferences and working groups
- Company outings, happy hours, hackathons, and tech talks
- Competitive compensation package with a strong benefits plan
- Remote work arrangement