Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Delinea

Senior Site Reliability Engineer – FedRAMP

Delinea

Senior Site Reliability Engineer owning reliability, monitoring, and incident response for Delinea’s FedRAMP SaaS identity security platform. Automating Azure and AWS operations.

Posted 9/10/2026full-timeRemote • 🇺🇸 United StatesSenior💰 $130,000 - $160,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in Site Reliability Engineering and DevOps, with a strong focus on managing production SaaS services, incident response, and infrastructure as code using Terraform and Azure DevOps. Proficient in monitoring and observability platforms like Datadog, with a solid understanding of cloud operations and security fundamentals.

Highest-signal resume keywords
Site Reliability EngineeringAzure ExperienceDatadog MonitoringInfrastructure As CodeIncident Response

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringAzureDatadogTerraformKubernetesPowerShellPythonCI/CDYAMLJSON
Soft Skills
Clear Written Communication
Tools & Technologies
Azure DevOpsWeb Application FirewallObservability PlatformMonitoring ToolsIncident Management Tools
Industry Keywords
SaaSCloud OperationsFedRAMP HighService-Level IndicatorsService-Level Objectives

Tech Stack

Tools & technologies
AzureCloudDNSFirewallsKubernetesPythonRedisSQLTerraform

About the role

Key responsibilities & impact
  • Own reliability for production SaaS services end to end, including availability, performance, and capacity
  • Define service-level indicators, service-level objectives, and error budgets
  • Build and tune monitoring in Datadog and Azure Monitor, including detection monitors, synthetic checks, dashboards, and alert routing
  • Automate incident response and replace manual runbook steps with code
  • Participate in on-call rotation and lead incident response for high-severity events
  • Coordinate incident resolution across Support, Engineering, and Product
  • Write post-incident reviews and customer-facing root cause analyses
  • Drive preventive actions to completion
  • Work support escalations by reproducing issues, diagnosing logs, traces, and network captures, and routing or resolving issues with evidence
  • Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines
  • Administer the web application firewall, including rule tuning, rate limiting, and false-positive triage
  • Manage observability platform costs, including Datadog indexing, retention, custom metrics, APM, log ingestion, and archived log storage
  • Operate within the FedRAMP High environment under change control and continuous monitoring processes
  • Improve on-call rotation design, escalation paths, alert quality, runbook coverage, and regional handoffs
  • Partner with Support, Security, Product, and Development to launch services with monitoring, runbooks, and SLOs

Requirements

What you’ll need
  • 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product
  • Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and cloud security fundamentals
  • Production experience with an observability platform such as Datadog, including metrics, logs, APM, dashboards, and monitor design
  • Demonstrated ownership of SLIs, SLOs, and error budgets
  • Incident response experience, including running incident bridges and writing postmortems
  • Kubernetes production experience, including ingress, deployments, resource limits, and troubleshooting failing workloads
  • Infrastructure as code with Terraform
  • CI/CD pipeline creation and troubleshooting; Azure DevOps preferred
  • Scripting in PowerShell and Python
  • Fluency with YAML and JSON
  • Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting
  • Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments
  • Clear written communication
  • Willingness to participate in an on-call rotation covering weekends and emergencies
  • Up to 10% travel
  • Prior government cloud experience is welcome but not required
  • Must not require any type of U.S. work authorization now or in the future

Benefits

Comp & perks
  • Equity
  • Performance-based bonus program or role-based incentive programs
  • Healthcare insurance
  • Pension/retirement matching
  • Comprehensive life insurance
  • Employee assistance program
  • Time off plans
  • Paid company holidays
  • Career progression
  • Meaningful work and culture of innovation