FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on backend infrastructure, monitoring, and automation. Proficient in optimizing cloud services and databases to enhance performance and reliability while effectively collaborating with cross-functional teams.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Google Cloud Platform (GCP)Monitoring and Alerting StacksScripting and Automation (Python, Bash)Database Performance Optimization (Postgres)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringBackend EngineeringMonitoring and AlertingScripting (Python, Bash)Database OptimizationCloud Infrastructure (GCP)Containerization (Docker)CI/CD PipelinesCapacity PlanningIncident Management
Soft Skills
Strong Communication SkillsCollaborationRelationship BuildingEscalation ManagementBlameless Postmortems
Tools & Technologies
GrafanaPrometheusZabbixDatadogCloud RunCloud SQLGoogle Cloud Storage (GCS)
Industry Keywords
Infrastructure OptimizationReliability MetricsOn-Call RotationML Training InfrastructureIncident Response
Tech Stack
Tools & technologiesCloudDockerGoogle Cloud PlatformGrafanaPostgresPrometheusPythonSQL
About the role
Key responsibilities & impact- Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on root causes.
- Level up our monitoring and alerting, and build out auto-remediation, so on-call load scales with automation, not headcount.
- Partner with our agentic engineering work to build agents that triage alerts and handle routine remediation.
- Harden and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance as load scales.
- Own database scale and performance; connection pooling, query optimization and indexing, read replicas, and capacity planning, so Postgres doesn't become the bottleneck as data volume grows.
- Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team.
- Run blameless postmortems and drive fixes for root causes, not just symptoms.
- Participate in on-call rotation.
Requirements
What you’ll need- 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership.
- Deep experience with a major cloud provider (GCP preferred); compute, managed databases, object storage, networking.
- Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar).
- Strong scripting/automation skills (Python, Bash, or similar).
- Comfortable with containerized workloads (Docker) and CI/CD pipelines.
- Track record of reducing incident volume or improving reliability metrics — not just responding to incidents.
- Strong communication skills, comfortable working with both technical and non-technical stakeholders, know when and how to escalate urgency, and build strong working relationships across teams.
- Strong communication skills in English — you write clearly and engage well async.
Benefits
Comp & perks- Health insurance
- 401(k) matching
- Flexible work hours
- Paid time off
