Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Outpost

Site Reliability Engineer

Outpost

Site Reliability Engineer ensuring system reliability for a logistics platform. Working on backend/API, monitoring, and proactive incident response with a high-performance team.

Posted 7/31/2026contractRemote • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on backend infrastructure, monitoring, and automation. Proficient in optimizing cloud services and databases to enhance performance and reliability while effectively collaborating with cross-functional teams.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Google Cloud Platform (GCP)Monitoring and Alerting StacksScripting and Automation (Python, Bash)Database Performance Optimization (Postgres)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringBackend EngineeringMonitoring and AlertingScripting (Python, Bash)Database OptimizationCloud Infrastructure (GCP)Containerization (Docker)CI/CD PipelinesCapacity PlanningIncident Management
Soft Skills
Strong Communication SkillsCollaborationRelationship BuildingEscalation ManagementBlameless Postmortems
Tools & Technologies
GrafanaPrometheusZabbixDatadogCloud RunCloud SQLGoogle Cloud Storage (GCS)
Industry Keywords
Infrastructure OptimizationReliability MetricsOn-Call RotationML Training InfrastructureIncident Response

Tech Stack

Tools & technologies
CloudDockerGoogle Cloud PlatformGrafanaPostgresPrometheusPythonSQL

About the role

Key responsibilities & impact
  • Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on root causes.
  • Level up our monitoring and alerting, and build out auto-remediation, so on-call load scales with automation, not headcount.
  • Partner with our agentic engineering work to build agents that triage alerts and handle routine remediation.
  • Harden and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance as load scales.
  • Own database scale and performance; connection pooling, query optimization and indexing, read replicas, and capacity planning, so Postgres doesn't become the bottleneck as data volume grows.
  • Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team.
  • Run blameless postmortems and drive fixes for root causes, not just symptoms.
  • Participate in on-call rotation.

Requirements

What you’ll need
  • 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership.
  • Deep experience with a major cloud provider (GCP preferred); compute, managed databases, object storage, networking.
  • Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar).
  • Strong scripting/automation skills (Python, Bash, or similar).
  • Comfortable with containerized workloads (Docker) and CI/CD pipelines.
  • Track record of reducing incident volume or improving reliability metrics — not just responding to incidents.
  • Strong communication skills, comfortable working with both technical and non-technical stakeholders, know when and how to escalate urgency, and build strong working relationships across teams.
  • Strong communication skills in English — you write clearly and engage well async.

Benefits

Comp & perks
  • Health insurance
  • 401(k) matching
  • Flexible work hours
  • Paid time off