Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Granicus

Site Reliability Engineer – Level 3

Granicus

Site Reliability Engineer 3 modernizing reliability engineering for Granicus with a focus on AIOps and automation. Improve service reliability and build scalable, resilient platforms for various workloads.

Posted 7/29/2026full-timeRemote • 🇮🇳 IndiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AIOps implementation, observability stack management, and incident management within large-scale cloud environments. Proficient in leveraging ELK/OpenSearch for log ingestion, monitoring, and alerting while driving improvements in system reliability and operational efficiency.

Highest-signal resume keywords
AIOps ImplementationELK/OpenSearch ExpertiseIncident ManagementCloud Platforms (AWS/Azure/GCP)Infrastructure as Code (Terraform/Ansible)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AIOps WorkflowsLog IngestionAdvanced Kibana QueryingTelemetry PreparationRCA (Root Cause Analysis)Anomaly DetectionAlert CorrelationDistributed SystemsPerformance TuningSLO-Based Reliability
Soft Skills
CollaborationProblem-SolvingDocumentationKnowledge Sharing
Tools & Technologies
ELK StackOpenSearchCloud PlatformsITSM/Ticketing SystemsChatOpsCMDB/Runbook RepositoriesAutomation Platforms
Certifications & Qualifications
AWS DevOps EngineerAWS ML SpecialtyGoogle Cloud DevOps EngineerAzure DevOps EngineerKubernetes/CKA
Industry Keywords
Production SupportObservabilityOperational Best PracticesCapacity PlanningService Mapping

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudDistributed SystemsElasticSearchGoogle Cloud PlatformITSMKubernetesLinuxLogstashTerraformUnix

About the role

Key responsibilities & impact
  • Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration.
  • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups.
  • Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to accelerate diagnosis while validating AI recommendations before action.
  • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting.
  • Design and maintain observability across logs, metrics, and traces, ensuring actionable monitoring and high signal-to-noise alerting.
  • Build and enhance workflows for alerting, anomaly detection, and incident enrichment to reduce noise and improve accuracy.
  • Implement AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment.
  • Develop automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback plans, and auditability.
  • Implement AIOps remediation patterns that connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps.
  • Drive improvements in system reliability, performance, scalability, and resilience through engineering-led initiatives.
  • Partner with engineering teams to improve deployment safety, operational readiness, and production stability.
  • Maintain high-quality runbooks, documentation, and knowledge bases to improve on-call effectiveness and knowledge sharing.
  • Support capacity planning, performance tuning, and SLO-based reliability practices.
  • Apply security, access control, and operational guardrails across systems and automation.
  • Own AIOps implementation from use-case definition through production rollout, including telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.

Requirements

What you’ll need
  • 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments.
  • Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP) .
  • Expert in ELK/OpenSearch, including: Log ingestion (Logstash / Beats) Elasticsearch index design, scaling, and tuning Advanced Kibana querying and debugging Dashboards, alerts, and observability patterns for production systems
  • Hands-on experience in logs, metrics, and tracing .
  • Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata.
  • Solid understanding of incident management, RCA, SLOs, and operational best practices .
  • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting .
  • Hands-on implementation experience with AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation.
  • Ability to implement AIOps integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.
  • Experience measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, faster MTTD/MTTR, RCA quality, automation adoption, repeat usage, and business impact linkage.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar.
  • Preferred Certifications : AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.

Benefits

Comp & perks
  • Employee Resource Groups to encourage diverse voices
  • Coffee with Mark sessions – Our employees get to interact with our CEO on very important and sometimes difficult issues ranging from mental health to work-life balance and current affairs.
  • Microsoft Teams communities focused on wellness, art, furbabies, family, parenting, and more.
  • Special guests from time to time to discuss issues that impact our employee population