Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Ontrac Solutions

Site Reliability Engineer

Ontrac Solutions

Site Reliability Engineer keeping user-facing services and production systems reliable for a client's cloud operations team. Automating infrastructure, improving monitoring, and scaling Kubernetes-based deployments.

Posted 8/9/2026contractRemote • 🇲🇬 MadagascarMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in cloud infrastructure management, automation, and monitoring, with a strong focus on security and scalability. Proficient in using tools like Ansible, Terraform, and Kubernetes to enhance deployment processes and maintain high availability.

Highest-signal resume keywords
Infrastructure Automation with AnsibleKubernetes Deployment on EKSMonitoring and Alerting with PrometheusStrong Programming Skills in PythonConfiguration Management with Terraform

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure AutomationCloud-Native DeploymentsMonitoring and AlertingDebugging Production IssuesConfiguration ManagementProgramming in PythonProgramming in JavaProgramming in GolangProgramming in Node.jsLinux and Windows Familiarity
Soft Skills
CollaborationAsynchronous CommunicationDocumentation SkillsProblem-Solving Attitude
Tools & Technologies
AnsiblePuppetTerraformKubernetesPrometheusNginxHAProxyDocker
Industry Keywords
Cloud InfrastructureSRE KPIsProduction AvailabilityIncident ResponseConfiguration-Management Systems

Tech Stack

Tools & technologies
AnsibleAWSCloudDockerGoHAProxyJavaJavaScriptKubernetesLinuxNGINXNode.jsPrometheusPuppetPythonTerraform

About the role

Key responsibilities & impact
  • Participate in an on-call rotation responding to production availability incidents and support service engineers with customer incidents
  • Use on-call shifts to prevent incidents from recurring
  • Operate infrastructure with Ansible, Puppet, Terraform, and Kubernetes
  • Build monitoring and alerting focused on symptoms rather than outages
  • Document actions and convert findings into repeatable procedures and automation
  • Improve the deployment process
  • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
  • Debug production issues across services and stack levels
  • Plan infrastructure growth
  • Code infrastructure automation with Ansible and Terraform
  • Improve Prometheus monitoring and build new metrics
  • Help release managers deploy and fix new application versions
  • Plan and execute migration from AWS virtual machines to cloud-native, container-based Kubernetes deployments on EKS
  • Develop relationships with product groups and define SRE KPIs

Requirements

What you’ll need
  • Think cloud-first across public cloud environments
  • Think security-first
  • Understand systems, edge cases, failure modes, behaviors, and specific implementations
  • Familiarity with Linux and Windows
  • Knowledge of configuration-management systems such as Ansible or Puppet
  • Strong programming skills in Python, Java, Golang, or Node.js
  • Ability to collaborate and communicate asynchronously and document work
  • Go-for-it attitude and willingness to fix broken systems
  • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies

Benefits

Comp & perks
  • 🌐 Worldwide ❌ Jobs You've Hidden ⭐️ Saved Jobs ✅ Applied Jobs ✉️ Email Alerts 👤 Account Ontrac Solutions Website LinkedIn All Job Openings 11 - 50 employees Founded 2010 🤖 Artificial Intelligence 💼 Consulting 🤝 B2B Artificial Intelligence
  • Consulting
  • B2B Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes. Site Reliability Engineer Job not on LinkedIn 🔥 0 minutes ago 🇲🇬 Madagascar – Remote ⏳ Contract/Temporary 🟡 Mid-level 🟠 Senior ⛑ DevOps & Site Reliability Engineer (SRE) Ansible AWS Cloud Docker HAProxy Java JavaScript Kubernetes Linux NGINX Node.js Prometheus Puppet Python Terraform Go Apply Now Customize resume + cover letter Report problem ☆ Save ☑️ Mark as applied ❌ Hide 📋 Description
  • Participate in an on-call rotation responding to production availability incidents and support service engineers with customer incidents
  • Use on-call shifts to prevent incidents from recurring
  • Operate infrastructure with Ansible, Puppet, Terraform, and Kubernetes
  • Build monitoring and alerting focused on symptoms rather than outages
  • Document actions and convert findings into repeatable procedures and automation
  • Improve the deployment process
  • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users
  • Debug production issues across services and stack levels
  • Plan infrastructure growth
  • Code infrastructure automation with Ansible and Terraform
  • Improve Prometheus monitoring and build new metrics
  • Help release managers deploy and fix new application versions
  • Plan and execute migration from AWS virtual machines to cloud-native, container-based Kubernetes deployments on EKS
  • Develop relationships with product groups and define SRE KPIs 🎯 Requirements
  • Think cloud-first across public cloud environments
  • Think security-first
  • Understand systems, edge cases, failure modes, behaviors, and specific implementations
  • Familiarity with Linux and Windows
  • Knowledge of configuration-management systems such as Ansible or Puppet
  • Strong programming skills in Python, Java, Golang, or Node.js
  • Ability to collaborate and communicate asynchronously and document work
  • Go-for-it attitude and willingness to fix broken systems
  • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies Apply Now 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score 🌐 Worldwide Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or support@remoterocketship.com Search Search Jobs by country Search jobs by city Search jobs by job title Search entry-level jobs Search junior-level jobs Search senior-level jobs Search jobs by tech stack Search jobs by contract type Search remote internships Search remote part-time jobs Remote jobs Anywhere in the World Companies Hiring Anywhere in the World Companies Hiring Sales People Anywhere in the World Companies Hiring Software Engineers Anywhere in the World Resources Advice Tips for finding remote jobs Interview questions and answers Resume examples Cover letter examples Post a job Affiliates About us Is Remote Rocketship legit? Privacy policy Terms of service Job board SEO course Remote Job Search MasterClass AI Apply Copilot OpenClaw job finder Find jobs using your resume Jobs by Country Remote jobs anywhere in the world (Worldwide remote jobs) Remote jobs United States Remote jobs Australia Remote jobs Brazil Remote jobs Canada Remote jobs France Remote jobs Ireland Remote jobs Germany Remote jobs Netherlands Remote jobs Spain Remote jobs UK Popular Jobs Remote data analyst jobs Remote customer support jobs Remote executive assistant jobs Remote marketing jobs Remote product designer jobs Remote product manager jobs Remote project manager jobs Remote recruiter jobs Remote sales jobs Remote software engineer jobs Jobs by Type Remote full-time jobs Remote part-time jobs Remote contract jobs Remote internship jobs Remote entry-level jobs Remote jobs with no experience required Remote junior jobs (1-3 years of experience) Digital nomad jobs Remote jobs with no degree required Freelance remote jobs Temporary remote jobs Remote jobs hiring now Stay at home mom jobs