Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Akamai Technologies

Site Reliability Engineer – Guardicore AI Platform

Akamai Technologies

Site Reliability Engineer responsible for the reliability of Akamai's Guardicore AI Platform. Collaborating across teams to enhance performance and conduct complex investigations.

Posted 7/23/2026full-timeRemote • 🇪🇸 SpainMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in operating and enhancing Kubernetes infrastructure, with a strong focus on observability, security, and performance. Proven ability to lead cross-team initiatives and implement automation strategies using AI tools.

Highest-signal resume keywords
KubernetesObservability StrategyTroubleshooting SkillsCI/CDPython

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesDockerHelmPythonGoBashGitOpsInfrastructure as CodePrometheusGrafana
Soft Skills
Technical LeadershipProblem-SolvingCollaboration
Tools & Technologies
GCPAzureLinodeAWS
Industry Keywords
SREDevOpsPlatform EngineeringMicroservicesData Pipelines

Tech Stack

Tools & technologies
AWSAzureDockerGoGoogle Cloud PlatformGrafanaKubernetesLinuxMicroservicesPrometheusPython

About the role

Key responsibilities & impact
  • Operating secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling.
  • Enhancing platform reliability, observability, security, performance, and cost efficiency.
  • Providing guidance to engineers and developers to increase confidence that their services are performing as expected.
  • Leading complex production investigations and driving long-term improvements.
  • Leverage LLMs and AI-driven automation to auto-remediate incidents and streamline operations.
  • Partner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problems.
  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues.

Requirements

What you’ll need
  • 3+ years of experience in SRE, DevOps, or Platform Engineering, with a proven track record of mastering and troubleshooting complex system architectures.
  • Demonstrate ability to design and implement a comprehensive monitoring and observability strategy using tools like Prometheus and Grafana.
  • Have production experience with Kubernetes, Docker, Helm, and third-party clouds (GCP, Azure, Linode, AWS) on Linux-based infrastructure.
  • Have exceptional troubleshooting and problem-solving skills across network, system, applications, and database layers.
  • Have experience with GitOps, CI/CD, and Infrastructure as Code.
  • Have scripting and programming proficiency in Python, Go, and Bash.
  • Leverage AI tools in daily operational tasks and actively propose initiatives to improve platform automation.
  • Demonstrate technical leadership and ownership in driving cross-team initiatives, defining tools, and building foundational frameworks.

Benefits

Comp & perks
  • We support your health, well-being, finances, and life beyond work.
  • FlexBase adapts to your job's needs
  • It's about supporting employees to do their best work. We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.