FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Reliability and Platform Engineer
Grupo PernambucanasSenior Reliability and Platform Engineer responsible for improving platform observability and automation solutions at Grupo Pernambucanas. Driving incident solutions and enhancing technical reliability.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) and DevOps practices, focusing on observability, incident management, and automation. Proficient in cloud technologies and infrastructure as code, with a strong emphasis on improving platform performance and resilience.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Cloud Technologies (GCP, AWS, Azure)Kubernetes (CKA/CKAD)Incident ManagementAIOps Solutions
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Observability ToolsAutomationScriptingCI/CDAPI ManagementRoot Cause AnalysisMetrics MonitoringIncident ResponsePerformance OptimizationDisaster Recovery
Tools & Technologies
DatadogKubernetesAIOpsAPI GatewaysMonitoring Tools
Certifications & Qualifications
Cloud Certifications (GCP, AWS, Azure)Kubernetes Certifications (CKA, CKAD)
Industry Keywords
Financial EnvironmentsRetail EnvironmentsOperational RunbooksPlaybooksAnomaly Detection
Tech Stack
Tools & technologiesAWSAzureCloudGoogle Cloud PlatformKubernetes
About the role
Key responsibilities & impact- Review and evolve alerts, monitors, and triggering criteria.
- Ensure proper escalation and reduce ignored or unowned alerts.
- Implement and monitor SLIs, SLOs, and availability metrics.
- Improve platform observability: logs, metrics, tracing, and APM.
- Automate incident responses, diagnostics, and recoveries.
- Create and maintain operational runbooks and playbooks.
- Drive technical actions resulting from post-mortems.
- Reduce recurring failures and operational manual work (toil).
- Support capacity, performance, resilience, and disaster recovery.
- Evolve internal platform components and patterns.
- Provide technical support to DevOps, Platform, and Development teams.
- Explore and apply AIOps solutions for anomaly detection and intelligent alert correlation.
Requirements
What you’ll need- Senior experience in SRE, DevOps, or Platform Engineering.
- Experience with Datadog or an equivalent observability tool.
- Experience with cloud, Kubernetes, and infrastructure as code.
- Knowledge of CI/CD, automation, and scripting or development.
- Experience with incident management and root cause analysis.
- Familiarity with APIs, API gateways, rate limiting, and distributed observability.
- Ability to develop automations and internal tools.
- Hands-on experience with AIOps: anomaly detection, automatic alert correlation, or AI-assisted root cause analysis (plus).
- Experience in high-volume financial or retail environments (plus).
- Cloud (GCP, AWS, or Azure) or Kubernetes (CKA/CKAD) certifications (plus).
Benefits
Comp & perks- Bradesco Saúde medical plan (cost sharing)
- Odontoseg dental plan (subscription-based)
- Partnership with a dental clinic
- Meal voucher or food allowance
- Life insurance
- Profit-sharing (PPR)
- Transportation voucher
- Bicycle parking with locker rooms
- 10% discount on Pernambucanas products
- TotalPass (wellness platform)
- Mãe Pernambucanas program with prenatal support
- Partnerships with SESC and educational institutions for undergraduate and graduate courses
- Ângela Social Assistance, providing support to women in situations of violence.