FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer – FedRAMP
DelineaSenior Site Reliability Engineer owning reliability, monitoring, and incident response for Delinea’s FedRAMP SaaS identity security platform. Automating Azure and AWS operations.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in Site Reliability Engineering and DevOps, with a strong focus on managing production SaaS services, incident response, and infrastructure as code using Terraform and Azure DevOps. Proficient in monitoring and observability platforms like Datadog, with a solid understanding of cloud operations and security fundamentals.
Highest-signal resume keywords
Site Reliability EngineeringAzure ExperienceDatadog MonitoringInfrastructure As CodeIncident Response
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringAzureDatadogTerraformKubernetesPowerShellPythonCI/CDYAMLJSON
Soft Skills
Clear Written Communication
Tools & Technologies
Azure DevOpsWeb Application FirewallObservability PlatformMonitoring ToolsIncident Management Tools
Industry Keywords
SaaSCloud OperationsFedRAMP HighService-Level IndicatorsService-Level Objectives
Tech Stack
Tools & technologiesAzureCloudDNSFirewallsKubernetesPythonRedisSQLTerraform
About the role
Key responsibilities & impact- Own reliability for production SaaS services end to end, including availability, performance, and capacity
- Define service-level indicators, service-level objectives, and error budgets
- Build and tune monitoring in Datadog and Azure Monitor, including detection monitors, synthetic checks, dashboards, and alert routing
- Automate incident response and replace manual runbook steps with code
- Participate in on-call rotation and lead incident response for high-severity events
- Coordinate incident resolution across Support, Engineering, and Product
- Write post-incident reviews and customer-facing root cause analyses
- Drive preventive actions to completion
- Work support escalations by reproducing issues, diagnosing logs, traces, and network captures, and routing or resolving issues with evidence
- Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines
- Administer the web application firewall, including rule tuning, rate limiting, and false-positive triage
- Manage observability platform costs, including Datadog indexing, retention, custom metrics, APM, log ingestion, and archived log storage
- Operate within the FedRAMP High environment under change control and continuous monitoring processes
- Improve on-call rotation design, escalation paths, alert quality, runbook coverage, and regional handoffs
- Partner with Support, Security, Product, and Development to launch services with monitoring, runbooks, and SLOs
Requirements
What you’ll need- 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product
- Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and cloud security fundamentals
- Production experience with an observability platform such as Datadog, including metrics, logs, APM, dashboards, and monitor design
- Demonstrated ownership of SLIs, SLOs, and error budgets
- Incident response experience, including running incident bridges and writing postmortems
- Kubernetes production experience, including ingress, deployments, resource limits, and troubleshooting failing workloads
- Infrastructure as code with Terraform
- CI/CD pipeline creation and troubleshooting; Azure DevOps preferred
- Scripting in PowerShell and Python
- Fluency with YAML and JSON
- Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting
- Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments
- Clear written communication
- Willingness to participate in an on-call rotation covering weekends and emergencies
- Up to 10% travel
- Prior government cloud experience is welcome but not required
- Must not require any type of U.S. work authorization now or in the future
Benefits
Comp & perks- Equity
- Performance-based bonus program or role-based incentive programs
- Healthcare insurance
- Pension/retirement matching
- Comprehensive life insurance
- Employee assistance program
- Time off plans
- Paid company holidays
- Career progression
- Meaningful work and culture of innovation