Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
PNC

Software Engineering Manager – Site Reliability

PNC

Software Engineering Manager leading SRE teams and production reliability for PNC’s banking technology platforms. Driving incident response, observability, automation, and resilient 24x7 operations.

Posted 9/7/2026full-timePittsburgh • Alabama, Arizona, Colorado, Ohio, Pennsylvania, Texas • 🇺🇸 United StatesMid-LevelSenior💰 $100,100 - $185,900 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering and DevOps practices, with a strong focus on incident management, operational leadership, and system reliability improvements. Proven ability to lead and develop teams in high-availability environments while ensuring compliance and governance.

Highest-signal resume keywords
Site Reliability EngineeringIncident ManagementDevOps Best PracticesMonitoring ToolsChange Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production SupportAutomationCapacity PlanningPerformance TuningRoot Cause AnalysisETL MonitoringDatabase ManagementCloud PlatformsLinuxWindows
Soft Skills
LeadershipCoachingCollaborationAccountabilityCommunication
Tools & Technologies
DynatraceBigPandaLogscaleMongoDBCassandraOracleSQLElasticsearchRedisKafka
Industry Keywords
High AvailabilityDisaster RecoveryOperational MaturityGovernanceRisk Compliance

Tech Stack

Tools & technologies
CassandraCloudElasticSearchETLKafkaLinuxMongoDBOracleRedisSQL

About the role

Key responsibilities & impact
  • Lead, coach, and develop SRE and related teams
  • Set goals, drive accountability, and foster ownership, innovation, continuous learning, and DevOps/SRE best practices
  • Partner with cross-functional stakeholders to align technology and business objectives
  • Provide after-hours operational leadership and participate in on-call leadership rotations
  • Lead end-to-end response and remediation for major P1/P2 incidents
  • Guide triage, diagnostics, troubleshooting, service restoration, stakeholder communication, and post-incident analysis
  • Serve as escalation point for complex production issues across applications, infrastructure, databases, middleware, and integrations
  • Oversee batch/ETL monitoring and recovery processes
  • Lead root cause analysis and permanent remediation for recurring incidents
  • Oversee change management, release execution, readiness validation, rollback strategies, risk assessments, CAB reviews, and post-implementation reviews
  • Build and optimize monitoring, dashboards, alerting, and observability frameworks using enterprise tools including Dynatrace, BigPanda, and Logscale
  • Lead high availability, disaster recovery, failover, continuity testing, MTTR, uptime, and reliability improvements
  • Guide capacity planning and performance tuning for scalable systems
  • Manage distributed teams supporting critical systems in a global 24x7 operation
  • Automate incident remediation, monitoring, alerting, deployment, and validation; standardize runbooks and automation frameworks
  • Ensure governance, risk, compliance, audit support, vulnerability remediation, access management, and data governance
  • Manage development projects, development teams, and application support functions
  • Oversee application programming and analysis projects, including development, installation, and maintenance
  • Monitor adherence to quality standards and analyze applications against business needs and specifications

Requirements

What you’ll need
  • 5+ years of related experience
  • 3+ years of management experience
  • Strong experience in Site Reliability Engineering, Production Support, or DevOps
  • Proven ability to lead teams in high-availability, enterprise environments
  • Deep understanding of incident, problem, and change management frameworks
  • Hands-on knowledge of monitoring tools, cloud/infrastructure platforms, and automation
  • Experience improving system reliability, observability, and operational maturity
  • Experience with OCP under infrastructure; Linux/Windows
  • Experience with MongoDB and Cassandra; Oracle and SQL
  • Working knowledge of Elasticsearch, Redis, MQ, and Kafka is a plus
  • University/college degree typically required; comparable combination of education, job-specific certifications, and experience may be considered in lieu of a degree
  • No required certifications
  • No required licenses
  • No required or preferred language assessments
  • Must work in one of the listed Technology Hub locations and attend the office weekly
  • PNC will not provide sponsorship for employment visas or participate in STEM OPT

Benefits

Comp & perks
  • Medical/prescription drug coverage with a Health Savings Account feature
  • Dental and vision options
  • Employee and spouse/child life insurance
  • Short- and long-term disability protection
  • 401(k) with PNC match
  • Pension and stock purchase plans
  • Dependent care reimbursement account
  • Back-up child/elder care
  • Adoption, surrogacy, and doula reimbursement
  • Educational assistance, including select programs fully paid
  • Robust wellness program with financial incentives
  • Maternity and/or parental leave
  • Up to 11 paid holidays each year
  • 9 occasional absence days each year
  • 15 to 25 vacation days each year, depending on career level and years of service