Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Akamai Technologies

Site Reliability Engineer – II

Akamai Technologies

Site Reliability Engineer ensuring reliability and availability of Akamai’s critical security products. Collaborating with technology like Kubernetes and Kafka to improve monitoring and tools.

Posted 7/9/2026full-timeRemote • 🇨🇷 Costa RicaJuniorMid-Level💰 CRC 16,932,850 - CRC 30,479,150 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in automation, system reliability, and incident response within large-scale distributed systems. Proficient in utilizing tools and practices for continuous improvement and operational excellence in a collaborative engineering environment.

Highest-signal resume keywords
KubernetesPythonTerraformPrometheusDevOps

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Linux AdministrationUnix AdministrationAutomation DevelopmentMonitoring ToolsConfiguration ManagementIncident ResponseCapacity PlanningAutoscaling ConfigurationWorkload SchedulingSLO Definition
Soft Skills
CollaborationMentorshipAccountability
Tools & Technologies
GrafanaSaltStackAnsibleDistributed Tracing
Industry Keywords
SRESysAdminContinuous ImprovementOperational ExcellenceAutomation

Tech Stack

Tools & technologies
AnsibleAWSAzureDistributed SystemsDockerGoGrafanaKubernetesLinuxPrometheusPythonSaltStackSQLTerraformUnix

About the role

Key responsibilities & impact
  • Create solutions to improve automation and efficiency for systems and teams.
  • Optimize workflows, infrastructure, and applications.
  • Collaborate on deployment, monitoring, and resolving incidents.
  • Focus on reliability, scalability, and efficiency through automation and resource optimization.
  • Promote continuous improvement and operational excellence across all systems.
  • Provide support and mentorship for other engineers within the department.
  • Develop and maintain automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency.
  • Improve system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform.
  • Participate in on-call rotations, guiding restoration and repair of service-impacting issues.
  • Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.

Requirements

What you’ll need
  • Have 2 years of relevant experience and a Bachelor's degree in Computer Science or its equivalent.
  • Possess experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems.
  • Demonstrate experience in Kubernetes and large-scale containerization systems.
  • Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible.
  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

Benefits

Comp & perks
  • We support your health, well-being, finances, and life beyond work.
  • Health insurance
  • 401K savings plan
  • Company holidays
  • Vacation (in the form of PTO)
  • Sick time
  • Family friendly benefits including parental leave
  • Employee assistance program focusing on mental and financial wellness