FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineer – II
Akamai TechnologiesSite Reliability Engineer ensuring reliability and availability of Akamai’s critical security products. Collaborating with technology like Kubernetes and Kafka to improve monitoring and tools.
Posted 7/9/2026full-timeRemote • 🇨🇷 Costa RicaJuniorMid-Level💰 CRC 16,932,850 - CRC 30,479,150 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in automation, system reliability, and incident response within large-scale distributed systems. Proficient in utilizing tools and practices for continuous improvement and operational excellence in a collaborative engineering environment.
Highest-signal resume keywords
KubernetesPythonTerraformPrometheusDevOps
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux AdministrationUnix AdministrationAutomation DevelopmentMonitoring ToolsConfiguration ManagementIncident ResponseCapacity PlanningAutoscaling ConfigurationWorkload SchedulingSLO Definition
Soft Skills
CollaborationMentorshipAccountability
Tools & Technologies
GrafanaSaltStackAnsibleDistributed Tracing
Industry Keywords
SRESysAdminContinuous ImprovementOperational ExcellenceAutomation
Tech Stack
Tools & technologiesAnsibleAWSAzureDistributed SystemsDockerGoGrafanaKubernetesLinuxPrometheusPythonSaltStackSQLTerraformUnix
About the role
Key responsibilities & impact- Create solutions to improve automation and efficiency for systems and teams.
- Optimize workflows, infrastructure, and applications.
- Collaborate on deployment, monitoring, and resolving incidents.
- Focus on reliability, scalability, and efficiency through automation and resource optimization.
- Promote continuous improvement and operational excellence across all systems.
- Provide support and mentorship for other engineers within the department.
- Develop and maintain automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency.
- Improve system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform.
- Participate in on-call rotations, guiding restoration and repair of service-impacting issues.
- Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
- Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
Requirements
What you’ll need- Have 2 years of relevant experience and a Bachelor's degree in Computer Science or its equivalent.
- Possess experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems.
- Demonstrate experience in Kubernetes and large-scale containerization systems.
- Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible.
- Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
- Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.
Benefits
Comp & perks- We support your health, well-being, finances, and life beyond work.
- Health insurance
- 401K savings plan
- Company holidays
- Vacation (in the form of PTO)
- Sick time
- Family friendly benefits including parental leave
- Employee assistance program focusing on mental and financial wellness