See all jobs on JobTailor
Search thousands of fresh jobs every day.
- Fresh listings
- Fast filters
- No subscription required

Site Reliability Engineer
Kraken Digital Asset ExchangeImplement data infrastructure solutions (self service) that support the needs of dozens of business units and hundreds of engineers; Utilize Infrastructure as Code (IaC) principles to design, provision, and manage both on-premises and clou…
Core Competencies
Role fitUse this summary to align your resume positioning with the role.
Candidates should emphasize their expertise in managing and optimizing infrastructure systems, particularly through the use of Infrastructure as Code (IaC) principles and automation tools. Proficiency in data systems management, containerization, and cloud infrastructure, along with strong problem-solving abilities, is essential for success in this role.
ATS Keywords
Tailor your resumeTip: use these terms in your resume and cover letter to boost ATS matches.
Tech Stack
Tools & technologiesAbout the role
Key responsibilities & impact- Collaborate closely with diverse cross-functional teams to conceive, execute, and oversee foundational infrastructure systems.
- Guarantee the availability, high performance, scalability, and cost efficiency of critical services and platforms.
- Participate in system monitoring, incident response, automation initiatives, and infrastructure improvements.
- Implement data infrastructure solutions that support the needs of business units and engineers.
- Utilize Infrastructure as Code (IaC) principles to design, provision, and manage infrastructure components using tools such as Terraform.
- Develop and maintain automation scripts to automate operational tasks and deployments.
- Enhance and manage CI/CD pipelines for consistent software deployments across the data infrastructure.
- Implement robust data monitoring and alerting solutions to proactively detect anomalies and performance issues.
- Manage and implement role-based access control (RBAC) and permissions for user groups and workflows across environments.
- Document architecture, processes, and best practices to enable knowledge sharing and support continuous improvement.
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Software Engineering, or a related field (or equivalent experience).
- Proven experience of 1+ year as a Site Reliability Engineer, Infrastructure/Platform/DevOps Engineer, Software Engineer, or similar roles.
- Ability to leverage AI tools and agents such as Claude and OpenAI to efficiently deliver business value.
- Solid understanding of bash/shell scripting and proficiency in at least one programming language (preferably Python, Golang, or Rust).
- Experience with containerization tools such as Docker or Podman.
- Strong problem-solving skills and the ability to troubleshoot complex systems.
- Experience managing and operating data systems such as Kafka, Redis, ElasticSearch, MariaDB, AirFlow, Debezium, ScyllaDB, TiDB, Hashicorp Vault.
- Experience managing self-hosted and SaaS platforms such as Splunk, VictoriaMetrics, Grafana, Cloudflare, Ingresses, Gitlab.
- Experience running Kubernetes as a Platform offering for engineering teams.
- Working experience in managing AWS infrastructure components.