Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Bank of America

Site Reliability Engineer

Bank of America

Site Reliability Engineer responsible for ensuring the reliability and performance of systems at Bank of America. Designing highly scalable systems and collaborating with cross-functional teams for improvements.

Posted 7/31/2026full-timeCharlotte • North Carolina, Texas • 🇺🇸 United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and implementing scalable systems, with strong proficiency in Linux/Unix environments and scripting languages. Capable of utilizing monitoring tools and cloud platforms to ensure system reliability and performance while collaborating effectively with cross-functional teams.

Highest-signal resume keywords
Scalable System DesignLinux/Unix ProficiencyScripting Languages (Python, Shell, Perl)Monitoring Tools (Prometheus, Grafana, ELK, Splunk)Cloud Platforms (AWS, Azure, Google Cloud)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Scalable System DesignLinux/Unix ProficiencyScripting Languages (Python, Shell, Perl)APM Tools (DynaTrace)Networking Principles (TCP/IP, HTTP, DNS)Containerization Technologies (Docker, Kubernetes)Monitoring Tools (Prometheus, Grafana, ELK, Splunk)Troubleshooting SkillsAutomation of TasksDocumentation Creation
Soft Skills
Problem-Solving SkillsCommunication SkillsCollaboration SkillsAttention to DetailAdaptability in Fast-Paced Environments
Tools & Technologies
DynaTraceAWSAzureGoogle CloudDockerKubernetesPrometheusGrafanaELK StackSplunk
Industry Keywords
Service Level Objectives (SLOs)Service Level Agreements (SLAs)Site Reliability EngineeringPerformance BottlenecksPost-Incident Analysis

Tech Stack

Tools & technologies
AWSAzureCloudDNSDockerGrafanaKubernetesLinuxPerlPrometheusPythonSplunkTCP/IPUnix

About the role

Key responsibilities & impact
  • Design and implement highly available and scalable systems, ensuring the reliability and performance of the company's website or application
  • Collaborate with cross-functional teams to define and establish service level objectives (SLOs) and service level agreements (SLAs) for critical systems
  • Monitor systems and applications, proactively identifying and resolving any performance bottlenecks or availability issues
  • Develop and maintain monitoring tools, alerts, and dashboards to provide visibility into system health and performance
  • Conduct post-incident analyses to identify root causes and implement preventive measures to avoid future incidents
  • Automate repetitive tasks and processes to improve efficiency and reduce manual intervention
  • Create and maintain documentation for system architecture, configuration, and troubleshooting procedures
  • Collaborate with development teams to implement and deploy new features and enhancements, ensuring they meet reliability and performance standards
  • Stay up to date with industry best practices, new technologies, and emerging trends in site reliability engineering

Requirements

What you’ll need
  • 5+ years of experience of designing and implementing scalable systems
  • Strong knowledge of Linux/Unix systems and command line tools
  • Proficiency in scripting languages such as Python, Shell, or Perl
  • Experience with APM tools such as DynaTrace
  • Familiarity with cloud platforms like AWS, Azure, or Google Cloud
  • Understanding of networking principles and protocols (TCP/IP, HTTP, DNS, etc.)
  • Knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools
  • Knowledge in monitoring and logging tools such as Prometheus, Grafana, ELK stack, or Splunk
  • Strong problem-solving and troubleshooting skills, with the ability to analyze and resolve complex technical issues
  • Excellent communication and collaboration skills to work effectively with cross-functional teams
  • Strong attention to detail and ability to work in a fast-paced, dynamic environment

Benefits

Comp & perks
  • Health insurance
  • Retirement plans
  • Paid time off
  • Flexible work arrangements
  • Professional development opportunities