Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CSC Generation

Site Reliability Engineer

CSC Generation

Site Reliability Engineer ensuring reliability, performance, and scalability across Backcountry's multi-cloud platform. Collaborating with engineering teams to improve systems and support features.

Posted 7/21/2026full-timeRemote • 🇨🇷 Costa RicaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in service resiliency, performance tuning, and system design, with a strong focus on automation and observability in cloud environments. Proficient in leveraging AI-assisted engineering tools and managing containerized production services across multi-cloud platforms.

Highest-signal resume keywords
Kubernetes ManagementInfrastructure as Code (Terraform, AWS CDK, Ansible)Cloud Experience (Google Cloud Platform, AWS)Observability Tooling (Grafana, Prometheus, Loki)Scripting and Programming (Bash, Python, TypeScript)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Kubernetes ManagementInfrastructure as CodeCloud ExperienceScripting Languages (Bash, Python, TypeScript)Observability ToolingAI-Assisted Coding ToolsLinux ManagementInternet Application ProtocolsGitOpsDevOps Practices
Soft Skills
Advanced English Communication
Tools & Technologies
Claude CodeGitHub CopilotTerraformAWS CDKAnsibleGrafanaPrometheusLokiOpenSearchArgoCD
Industry Keywords
Site Reliability EngineeringService Level Indicators (SLIs)Service Level Objectives (SLOs)FinOpsMulti-Cloud Stack

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudDNSGoogle Cloud PlatformGrafanaJavaScriptKubernetesLinuxNode.jsPrometheusPythonTerraformTypeScript

About the role

Key responsibilities & impact
  • Work on service resiliency, performance tuning, and system design across Backcountry's platform
  • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
  • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories
  • Reduce toil by designing and implementing automation
  • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads
  • Monitor system health and capacity, taking proactive action to fix problems before they occur
  • Collaborate with engineering teams to build, deploy, and support features
  • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services
  • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy
  • Participate in the on-call support rotation within the SRE team

Requirements

What you’ll need
  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written.

Benefits

Comp & perks
  • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.