Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Bank of America

Site Reliability Engineer – Lead

Bank of America

Site Reliability Engineer Lead role at Bank of America driving reliability and performance across technology platforms. Leading a team focused on modern Agile practices in large-scale environments.

Posted 7/24/2026full-timePlano • Arizona, North Carolina, Texas • 🇺🇸 United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in Site Reliability Engineering (SRE) and platform engineering, with a strong focus on developing technology strategies, implementing observability practices, and leading incident response efforts. Proven ability to mentor teams and drive automation in large-scale environments while ensuring compliance with industry standards.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure-as-CodeObservability StacksCloud PlatformsIncident Response Leadership

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Systems EngineeringDevOpsLinux/Unix SystemsWindows SystemsNetworkingDistributed ComputingCapacity PlanningPerformance TuningCost OptimizationAutomation Tools
Soft Skills
Excellent CommunicationStakeholder ManagementExecutive Engagement
Tools & Technologies
TerraformAnsiblePythonKubernetesDynatraceGrafanaSplunkOpenTelemetry
Industry Keywords
Agile Solution DeliveryGovernance ModelsProactive MonitoringRoot Cause AnalysisTelemetry

Tech Stack

Tools & technologies
AnsibleCloudGrafanaKubernetesLinuxPythonSplunkTerraformUnix

About the role

Key responsibilities & impact
  • Build and lead a team to deliver technology products and services
  • Develop a technology strategy and ensure technology solutions comply with standards
  • Promote design, engineering, and organizational practices
  • Advocate and advance modern, Agile solution delivery practices
  • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols
  • Establish governance models for reliability engineering across distributed teams
  • Champion a culture of observability and proactive monitoring
  • Lead root cause analysis (RCA) and post-incident reviews
  • Implement proactive problem detection using telemetry
  • Develop and maintain capacity models and monitor performance trends
  • Drive automation of operational tasks including deployments and scaling
  • Oversee major incident response and communication processes
  • Serve as a senior technical advisor and thought leader in SRE and platform engineering
  • Mentor SRE teams and partner with engineering leaders across the enterprise

Requirements

What you’ll need
  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry)
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python)
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes)
  • Demonstrated success in leading incident response and driving systemic improvements
  • Experience with capacity planning, performance tuning, and cost optimization
  • Excellent communication and stakeholder management skills, including executive engagement.

Benefits

Comp & perks
  • Health insurance
  • Retirement plans
  • Paid time off
  • Flexible work arrangements
  • Professional development programs