Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Arista Networks

Senior Site Reliability Engineer – CloudVision

Arista Networks

Site Reliability Engineer at Arista Networks focusing on building and operating critical production systems for scalability and reliability. Engaging in a collaborative remote role from Ireland.

Posted 8/3/2026full-timeRemote • 🇮🇪 IrelandSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing, building, and deploying scalable and reliable production systems while ensuring security compliance. Proficient in automation, incident response, and monitoring practices to enhance operational efficiency and system performance.

Highest-signal resume keywords
Infrastructure-As-Code PrinciplesContainer Orchestration (Kubernetes)Monitoring Stack Management (Prometheus, Grafana)Programming Languages (Go, Python, Bash)Cloud Platforms (GCP, AWS, Azure)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Linux AdministrationSoftware TroubleshootingServer ProvisioningDatabase Management (PostgreSQL)Performance TuningAutomation WorkflowsIncident ResponseContinuous Improvement MethodologiesVirtualization TechnologiesCI/CD Systems (GitLab, Spinnaker)
Soft Skills
Analytical Problem-SolvingCollaborative CommunicationDecisivenessMethodical ApproachStakeholder Engagement
Tools & Technologies
DockerTerraformArtifact RepositoriesDocker RegistriesMonitoring Infrastructure
Industry Keywords
ScalabilityReliabilityObservabilitySecurity StandardsDistributed Systems Architecture

Tech Stack

Tools & technologies
AWSAzureCloudDistributed SystemsDockerGoGoogle Cloud PlatformGrafanaKubernetesLinuxPostgresPrometheusPythonShell ScriptingSpinnakerTerraformUnix

About the role

Key responsibilities & impact
  • Design, build, and deploy production systems with a focus on scalability, reliability, observability, and performance, ensuring systems meet stringent security standards
  • Develop and maintain comprehensive automation solutions to eliminate toil and streamline operational efficiency across production environments
  • Proactively monitor production systems, establish intelligent alerting strategies, and implement automated incident response mechanisms to minimise downtime
  • Create and maintain detailed incident response runbooks; conduct thorough postmortem analyses following incidents to identify root causes and prevent recurrence
  • Collaborate with software engineering teams to identify and resolve infrastructural bottlenecks, designing innovative solutions that enhance product deployment workflows
  • Manage and optimise monitoring infrastructure using industry-standard tools, ensuring comprehensive visibility across all systems
  • Plan, communicate, and execute maintenance windows on production systems with minimal disruption to service availability
  • Triage platform and infrastructural issues with decisiveness and analytical rigour; engage with third-party vendors and support teams as required
  • Deploy new systems and updates in a staged, risk-managed manner, ensuring safe and incremental rollouts
  • Survey and adopt best practices in infrastructure and platform management to maintain secure, scalable, and fault-tolerant systems
  • Study the design and implementation details of open-source systems to enhance troubleshooting capabilities and accelerate issue resolution
  • Work transparently with stakeholders to communicate system status, planned maintenance, and infrastructure improvements

Requirements

What you’ll need
  • Bachelor's degree in Computer Science, Engineering, or equivalent professional experience (5+ years in a related infrastructure or systems role)
  • Proficiency in one or more programming languages: Go, Python, or bash shell scripting, with the ability to implement medium-complexity automation workflows
  • Strong knowledge of Linux or UNIX from both administration and debugging perspectives
  • Hands-on experience operating software systems, infrastructure, and complex applications at scale in production environments
  • Demonstrated expertise in infrastructure-as-code principles and practices
  • Strong problem-solving and software troubleshooting skills with a methodical, analytical approach
  • Experience with server provisioning, particularly from storage and networking perspectives
  • Proven ability to work collaboratively within cross-functional teams and communicate technical concepts clearly
  • Experience with incident response, postmortem analysis, and continuous improvement methodologies
  • Experience with container orchestration platforms, particularly Kubernetes
  • Hands-on experience with Docker and virtualisation technologies
  • Proficiency in managing monitoring stacks, including Prometheus and Grafana
  • Experience with CI/CD systems such as GitLab tools or Spinnaker
  • Knowledge of infrastructure-as-code frameworks, particularly Terraform
  • Experience managing databases such as PostgreSQL or equivalent relational database management systems
  • Experience with artifact repositories and Docker registries
  • Familiarity with cloud platforms (Google Cloud Platform, Amazon Web Services, or Microsoft Azure)
  • Understanding of distributed systems architecture and principles
  • Experience with performance tuning and system optimisation
  • Knowledge of security best practices in infrastructure and systems design
  • On-call support experience and comfort with incident response responsibilities

Benefits

Comp & perks
  • Professional development opportunities
  • Remote work options