Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Salesforce

Director, Site Reliability Engineering

Salesforce

Salesforce Director leading SRE strategy, observability, and automation for its AI-powered CRM platform. Transforming reactive operations into proactive, reliable, data-driven engineering.

Posted 8/4/2026full-timeNew York City • California, New York, Texas • 🇺🇸 United StatesLead💰 $197,300 - $313,700 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering, focusing on automation, observability, and operational excellence. Proven ability to lead high-performing teams and drive engineering transformations that enhance system reliability and performance.

Highest-signal resume keywords
Site Reliability Engineering StrategyObservability PracticesDistributed Systems UnderstandingEngineering LeadershipAutomation and AI Integration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringDistributed SystemsCloud ArchitectureIncident ManagementService-Level ObjectivesAutomationMetrics and Logs AnalysisRoot-Cause AnalysisPerformance EngineeringCapacity Planning
Soft Skills
Executive CommunicationCross-Functional CollaborationTeam LeadershipProblem-SolvingStrategic Thinking
Tools & Technologies
AWSNew RelicSplunkDatadogSentryHoneycombGrafanaPrometheusOpenTelemetryChaos Engineering
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree or MBA (Preferred)
Industry Keywords
Operational RiskResilience TestingDisaster RecoverySelf-Service CapabilitiesAI-Assisted Operations

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsGrafanaPrometheusSplunk

About the role

Key responsibilities & impact
  • Define and execute the long-term Site Reliability Engineering strategy and roadmap
  • Establish the SRE operating model, team scope, engagement models, ownership boundaries, and success measures
  • Build and develop a high-performing team of site reliability and operations engineers
  • Modernize SRE through automation, AI-assisted operations, self-service capabilities, and engineering-first practices
  • Translate business priorities and customer impact into reliability investments and engineering outcomes
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems
  • Define production operational and observability readiness, including readiness reviews and certification practices
  • Drive improvements in availability, performance, resiliency, and recovery
  • Define enterprise observability strategy across metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry
  • Establish instrumentation, telemetry, dashboards, alerting, and service-health standards
  • Improve incident detection, response, mitigation, communication, learning, and on-call practices
  • Lead automated detection, diagnosis, remediation, and incident creation
  • Establish blameless post-incident reviews and eliminate recurring operational toil
  • Develop intelligent-operations roadmaps covering anomaly detection, event correlation, automated triage, root-cause analysis, and remediation
  • Evaluate AI-assisted workflows across observability, incident response, capacity planning, and operational support
  • Build safe, measurable, auditable automation with appropriate human oversight
  • Partner with engineering, DevOps, and business stakeholders to build shared accountability for production outcomes
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination
  • Communicate reliability posture, risks, trends, and investments to executive and technical audiences

Requirements

What you’ll need
  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field
  • 10+ years of progressive engineering experience
  • 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
  • Proven experience building or transforming a reliability or operational engineering organization
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices
  • Demonstrated ability to improve reliability through engineering and automation
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems
  • Strong understanding of metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring
  • Demonstrated success driving alignment across cross-functional engineering teams and executive stakeholders
  • Ability to balance immediate operational needs with long-term engineering transformation
  • Strong written, verbal, and executive communication skills
  • Preferred: Master's degree or MBA
  • Preferred: Experience operating large-scale systems in AWS or another major cloud environment
  • Preferred: Experience with New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry
  • Preferred: Experience implementing OpenTelemetry or common instrumentation standards
  • Preferred: Experience building internal developer platforms, paved roads, or self-service reliability capabilities
  • Preferred: Experience applying AI, machine learning, or agent-based automation to operational workflows
  • Preferred: Experience with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering
  • Preferred: Software engineering experience and ability to engage in architecture and design discussions
  • Preferred: Experience supporting high-profile launches, events, or systems with significant customer and business impact

Benefits

Comp & perks
  • Time off programs
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Mental health support
  • Paid parental leave
  • Life insurance
  • Disability insurance
  • 401(k)
  • Employee stock purchasing program
  • Potential eligibility for incentive compensation
  • Potential eligibility for equity
  • Potential eligibility for benefits
  • Reasonable accommodation during the application or recruiting process