FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer
Akamai TechnologiesSenior Site Reliability Engineer building observability, automation, and Kubernetes tooling for Akamai's distributed cloud and edge platform. Improving reliability, scalability, and performance across Compute services.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering with a strong focus on Linux system administration, containerized environments, and automation. Proficient in deploying and maintaining observability platforms while ensuring system reliability and performance across teams.
Highest-signal resume keywords
Site Reliability EngineeringLinux System AdministrationKubernetes ManagementCI/CD PracticesAutomation/Scripting with Python
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringLinux System AdministrationKubernetesCI/CDAutomation/ScriptingNetworking FundamentalsInfrastructure as CodeConfiguration ManagementContainerized EnvironmentsTroubleshooting
Soft Skills
CollaborationProblem SolvingProactive Troubleshooting
Tools & Technologies
JenkinsGitPrometheusGrafanaTerraformAnsibleSaltStack
Certifications & Qualifications
Bachelor's Degree in Computer Science or Engineering
Industry Keywords
ObservabilityService ReliabilityScalabilityUsabilityCloud Storage Systems
Tech Stack
Tools & technologiesAnsibleCloudDNSGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRustSaltStackSwitchingTCP/IPTerraform
About the role
Key responsibilities & impact- Design, develop, and manage applications and infrastructure supporting Akamai's Compute products and services
- Create solutions that improve observability and enforce SLAs across internal teams
- Collaborate with operations and application development teams
- Create tooling and software that monitors and improves system reliability
- Solve complex problems through proactive troubleshooting, automation, and systems programming
- Deploy and maintain the observability platform and internal tooling
- Partner across teams to ensure product and service reliability, scalability, and usability
- Guide engineers and developers in assessing service performance
- Collaborate with support, operations, and engineering teams to investigate and troubleshoot complex problems
- Release new applications and modernize existing tooling
Requirements
What you’ll need- Bachelor's degree in Computer Science or Engineering
- 6 years of experience in Site Reliability Engineering or a related engineering role
- Linux system administration expertise
- Understanding of networking fundamentals, including TCP/IP, DNS, routing/switching, and storage concepts
- Hands-on experience with containerized environments and Kubernetes, including operating and troubleshooting production systems
- Knowledge of CI/CD and DevOps practices
- Hands-on experience using Jenkins, Git, Prometheus, and Grafana
- Experience with Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or similar tools
- Exposure to cloud storage systems
- Automation/scripting skills using Python, Bash, Go, Rust, or similar languages
Benefits
Comp & perks- Health, well-being, financial, and broader life-support benefits
- FlexBase flexible workplace options: work from home, in an office, or a combination of both