Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Jalasoft

Site Reliability Engineer – Azure, Observability, Scripting

Jalasoft

Site Reliability Engineer ensuring Jalasoft’s Azure and Kubernetes platforms remain reliable, scalable, and observable. Automating infrastructure, monitoring production systems, and strengthening incident and disaster recovery practices across Latin America.

Posted 8/5/2026full-timeRemote • 🇨🇴 ColombiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in cloud-native platform reliability and performance, particularly with Microsoft Azure and Kubernetes. Proficient in designing observability solutions, incident response processes, and disaster recovery strategies while ensuring operational excellence through automation and scripting.

Highest-signal resume keywords
Kubernetes OperationsAzure MonitorPrometheusGrafanaInfrastructure As Code

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesAzureScriptingTerraformBicepIncident ResponseDisaster RecoveryService Level IndicatorsMonitoringAutomation
Tools & Technologies
Azure DevOpsLog AnalyticsKQLMetricsExporters
Industry Keywords
Site Reliability EngineeringProduction OperationsCloud-Native PlatformsRPORTO

Tech Stack

Tools & technologies
AzureCloudGrafanaKubernetesPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes
  • Improve system availability, monitoring, incident response, and infrastructure automation in production environments
  • Design and implement observability solutions, including monitoring, metrics, exporters, alerting rules, and dashboards
  • Define and implement service level indicators, objectives, and error budgets
  • Design alerting and incident response processes, author runbooks, and support on-call practices
  • Design and test backup, restore, and disaster recovery solutions against RPO and RTO targets
  • Troubleshoot Kubernetes workloads, manage resources, and perform cluster upgrades
  • Read and modify infrastructure as code and Azure DevOps pipelines
  • Support operational excellence through automation and scripting

Requirements

What you’ll need
  • 6+ years of experience
  • 3+ years of experience operating Kubernetes in production
  • Site reliability engineering or production operations for Kubernetes workloads at scale
  • Experience with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy
  • Experience with Prometheus and Grafana, including metrics, exporters, recording and alerting rules, and dashboard design
  • Experience defining and implementing service level indicators, objectives, and error budgets
  • Experience designing alerting and incident response, including runbook authoring and on-call practice
  • Experience designing and testing backup, restore, and disaster recovery, including validation against RPO and RTO targets
  • Kubernetes operations experience, including workload troubleshooting, resource management, and cluster upgrades
  • Ability to read and modify infrastructure as code using Terraform or Bicep and Azure DevOps pipelines
  • Scripting experience in Python, PowerShell, or Bash
  • Professional working English

Benefits

Comp & perks
  • Remote work
  • 13 floating holiday
  • 15 vacation days per year completed
  • Good working environment
  • Equal opportunity employment without distinction based on protected characteristics