Dropbox

Site Reliability Engineer

Dropbox

full-time

Posted on:

Location Type: Remote

Location: Mexico

Visit company website

Explore more

AI Apply
Apply

About the role

  • Ensure the reliability, scalability, and performance of Dropbox's infrastructure and services
  • Collaborate with cross-functional teams to develop and maintain best practices for monitoring, logging, and incident response
  • Build, Implement and maintain automations & infrastructure-as-code tooling, specifically Terraform, Ansible, and Github Actions as well as custom code platforms
  • Utilize container orchestration platforms, such as Kubernetes, Amazon ECS and Red Hat Openshift, to manage containers at scale
  • Manage and optimize monitoring and logging pipelines using tools like Datadog and Cribl LogStream
  • Drive improvement projects related to service health and visibility for our stakeholders, ranging from developers to business service owners to C-level
  • Develop and maintain custom tooling and automation scripts in Bash, Python and other scripting languages

Requirements

  • 5+ years of experience in site reliability engineering or a similar engineering roles with hands-on coding experience
  • Strong knowledge of AWS services, including EC2, S3, RDS, R53, Lambda, and others
  • Strong knowledge of Linux administration, internals, filesystems, volume management and specific distro's such as Ubuntu, RHEL, DNS, DHCP
  • Experience with monitoring and logging tools, Datadog and logging pipeline tools such as Vector or Cribl LogStream
  • Experience driving one or more transformational programs related to metrics and observability
  • Experience with scripting in a higher level language (Python preferred)
  • Experience developing automation to solve infrastructure-related tasks with tools such as Chef/Ansible/Terraform
  • Experience with log analysis and building metrics, alerts and visuals from log data
  • Strong proficiency in infrastructure-as-code tools, such as Terraform
  • Strong Proficiency in Config Management tools specifically Ansible Automation Platform and Chef
  • Experience with containerization technologies, such as Docker, and container orchestration platforms like Kubernetes or Amazon ECS
  • Knowledge of LDAP, REST API's and current Auth
  • Familiarity with GitHub and Git-based workflows
  • Understanding of RDS databases and network security technologies, such as WAF
  • Strong problem-solving skills and the ability to work well in a fast-paced, collaborative environment
  • Excellent written and verbal communication skills.
Benefits
  • Competitive salary
  • Flexible work hours
  • Professional development budget
  • Home office setup allowance
  • Global team events
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills & Tools
site reliability engineeringinfrastructure-as-codescriptingLinux administrationmonitoring and loggingautomationcontainer orchestrationlog analysisconfig managementproblem-solving
Soft Skills
collaborationcommunicationorganizational skillsfast-paced workstakeholder engagement