Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Grupo Pernambucanas

DevOps Coordinator – SRE

Grupo Pernambucanas

Liderança técnica em DevOps & SRE no Grupo Pernambucanas. Responsável por operações de infraestrutura em ambientes financeiros e de e-commerce.

Posted 7/2/2026full-timeSão Paulo • 🇧🇷 BrazilMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing GCP and GKE production environments, ensuring scalability, health, and security while implementing CI/CD practices with tools like ArgoCD and Terraform. Proficient in incident management and observability using Datadog and Grafana, with a strong foundation in DevSecOps principles.

Highest-signal resume keywords
GCP & GKE Production Cluster OperationAdvanced Kubernetes ManagementCI/CD and GitOps ToolsObservability and SRE ConceptsDevSecOps and Incident Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Python ScriptingShell ScriptingTerraform ModulesHelm Charts CreationOpenTelemetry InstrumentationAPI Gateway ManagementSAST Tools (SonarQube)Cloud BuildArgoCDMulticloud Services (AWS, Azure)
Soft Skills
MentorshipTeam CollaborationCrisis Management
Tools & Technologies
DatadogGrafanaGoogle Kubernetes Engine (GKE)AWS Services (EKS, S3, IAM)Azure Services (AKS, Azure AD)
Certifications & Qualifications
CKAGoogle Cloud Professional Cloud ArchitectAWS Solutions ArchitectAZ-104
Industry Keywords
High-Availability EnvironmentsFinancial SystemsHigh-Volume E-CommerceCritical Transactions

Tech Stack

Tools & technologies
AWSAzureCloudGoogle Cloud PlatformGrafanaKubernetesNode.jsPythonTerraform

About the role

Key responsibilities & impact
  • Define and monitor SLOs with product squads, actively using error budgets to guide prioritization decisions and control release velocity.
  • Lead war rooms for critical incidents (P1/P2) end-to-end (triage, diagnosis, resolution) and conduct blameless post-mortems (5 Whys).
  • Operate and ensure the scalability, health, and security of production GKE (Google Kubernetes Engine) clusters, while maintaining visibility over workloads in AWS and Azure.
  • Design and evolve Cloud Build + ArgoCD pipelines with mandatory quality gates (SonarQube, image scanning, smoke tests) and define rollout strategies (canary, blue/green).
  • Structure and maintain Terraform modules for multi-project GCP environments, managing remote state, drift detection, and policy-as-code.
  • Ensure production readiness with structured logs, traces via OpenTelemetry, alerts in Datadog/Grafana, and integrate security tools (SAST, Trivy, Secret Manager) without introducing friction into the delivery flow.
  • Provide structured mentorship to the team (internal staff and consultants), use code reviews as a teaching tool, and organize a sustainable on-call rota.

Requirements

What you’ll need
  • GCP & GKE: Hands-on production cluster operation (troubleshooting, HPA, PDB, networking, node pools, Workload Identity).
  • Advanced Kubernetes: Deep understanding of workload lifecycle, RBAC, Network Policies, and Admission Controllers.
  • CI/CD, GitOps and IaC tools: ArgoCD, Cloud Build (or GitHub Actions/GitLab CI), and Terraform (modules with semantic versioning and state management).
  • Observability and SRE: Datadog and Grafana (dashboards, monitors, SLO tracking) and conceptual and practical mastery of SRE concepts (SLI/SLO/SLA).
  • DevSecOps & Incident Management: Experience with SAST (SonarQube), image scanning and crisis/incident handling.
  • Scripting: Python and/or Shell scripting for automation.
  • Multicloud knowledge of equivalent services in AWS (EKS, S3, IAM) and Azure (AKS, Azure AD).
  • Creation and maintenance of Helm charts.
  • Instrumenting services with OpenTelemetry.
  • Experience with API Gateways (Apigee X/Edge and/or Kong).
  • Active certifications: CKA, Google Cloud Professional Cloud Architect, AWS Solutions Architect or AZ-104.
  • Education: Bachelor's degree in Computer Science, Software Engineering, Information Systems or related fields (postgraduate degree is a plus).
  • Experience: Minimum of 6 years in DevOps/SRE/Infrastructure, with at least 2 years in coordination or technical leadership roles.
  • Industry experience: Prior work in high-availability environments, dealing with financial systems or high-volume e-commerce (critical transactions).
  • Language: Intermediate to advanced English for reading technical documentation and interacting with vendors.

Benefits

Comp & perks
  • Medical insurance (Bradesco Saúde) with co-payment
  • Dental insurance (Odontoseg) — optional enrollment
  • Partnership with a dental clinic
  • Meal allowance or grocery voucher
  • Life insurance
  • Profit-sharing (PPR)
  • Transportation allowance
  • Bicycle parking with changing rooms
  • 10% employee discount on Pernambucanas products
  • Partnerships with SESC and educational institutions for undergraduate and postgraduate courses