Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Valtech

Senior Site Reliability Engineer

Valtech

Senior Site Reliability Engineer at Valtech responsible for bridging software development and operations. Collaborating with multidisciplinary teams to enhance infrastructure and customer experience.

Posted 7/22/2026full-timeRemote • 🇵🇹 PortugalSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering with a strong focus on incident management, observability, and system architecture. Proficient in programming, scripting, and utilizing various monitoring and pipelining tools within cloud environments.

Highest-signal resume keywords
Site Reliability EngineeringIncident ManagementMonitoring SystemsCloud Services (AWS, Azure, GCP)Microservices Technology (Docker, Kubernetes)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringDevOps EngineeringQA EngineeringProgrammingScripting
Soft Skills
Assertive CommunicationLeadershipCoaching
Tools & Technologies
DatadogNew RelicDynatracePrometheusGrafanaGitHubAzure DevOpsGitLabJenkinsArgo CI/CD
Industry Keywords
Public-Facing Online ServiceECommerce PlatformsCorporate EnvironmentsInternational Context

Tech Stack

Tools & technologies
AWSAzureCloudDockerGoogle Cloud PlatformGrafanaJavaJenkinsKafkaKubernetesMicroservicesPrometheusSpring BootSpringBoot

About the role

Key responsibilities & impact
  • Work with teams to define SLIs and SLOs
  • Create systems for observability
  • Work with teams to analyze failure scenarios and possible mitigations
  • Create runbooks to remediate or prevent failure scenarios
  • Reduce work that does not add value
  • Participate and facilitate incident management, including on-call duty

Requirements

What you’ll need
  • 5 years of experience in the field of software engineering, DevOps engineering, QA engineering and/or cloud engineering
  • At least the last 2 years as a dedicated Site Reliability Engineer
  • Assertive with good communicative skills, capable of taking the lead and coaching a development team
  • Experience with incident management in a production environment of a public-facing online service
  • Experience in working in corporate environments
  • Experience programming and scripting
  • Basic knowledge of serverless services in one or more public cloud providers (AWS, Azure, GCP)
  • Extensive knowledge of and experience with various monitoring systems, APM systems such as Datadog, New Relic, Dynatrace, Prometheus, and Grafana
  • Knowledge of and experience with various pipelining tools, such as GitHub, Azure DevOps, GitLab, Jenkins
  • Knowledge of and experience with microservices-related technology: Docker, Kubernetes
  • Good conceptual understanding of software architecture and system thinking
  • Worked as an engineer in a DevOps context
  • Excellent command of English (C1 or above)
  • Familiar with Datadog, Argo CI/CD, Java / Springboot, Kafka, Kubernetes / EKS, AWS
  • Worked within the context of publicly accessible, highly available eCommerce platforms
  • Experience working in an international context with on- and off-shore teams

Benefits

Comp & perks
  • Flexibility, with remote and hybrid work options (country-dependent)
  • Career advancement, with international mobility and professional development programs
  • Learning and development, with access to cutting-edge tools, training and industry experts