FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering with a focus on .NET/C# applications, Windows Server environments, and AWS infrastructure. Proficient in incident response, monitoring, and automation to enhance system reliability and operational efficiency.
Highest-signal resume keywords
Site Reliability EngineeringDebugging .NET/C# ApplicationsPowerShell ScriptingGrafana and PrometheusAWS Infrastructure
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringDebugging .NET/C# ApplicationsPowerShell ScriptingPythonBashGrafanaPrometheusTerraformAurora PostgreSQLSLIs and SLOs
Soft Skills
Incident ResponseRoot Cause Analysis
Tools & Technologies
Windows ServerCI/CD PipelinesInfrastructure as Code (IaC)
Industry Keywords
Deployment StrategiesRelease AutomationRollback MechanismsOperational DocumentationMonitoring and Observability
Tech Stack
Tools & technologiesAWSGrafana.NETPostgresPrometheusPythonTerraform
About the role
Key responsibilities & impact- Own the day-to-day reliability of our .NET/C# services running on Windows.
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
- Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
- Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
- Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure.
- Improve deployment safety, release automation, and rollback strategies.
- Partner with developers to improve application operability, resilience, and fault isolation.
- Automate operational tasks through scripting and infrastructure automation.
- Create and maintain runbooks, operational documentation, and incident response procedures.
- Continuously improve monitoring, alert quality, automation, and platform reliability.
Requirements
What you’ll need- 3–5 years of experience in Site Reliability Engineering or related field
- Strong experience debugging and supporting .NET/C# applications in production.
- Hands-on experience with Windows Server environments.
- Strong PowerShell scripting skills.
- Experience with Python or Bash.
- Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
- Experience with modern CI/CD pipelines.
- Knowledge of deployment strategies, release automation, and rollback mechanisms.
- Experience working with AWS.
- Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
- Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
- Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.
Benefits
Comp & perks- Work on mission-critical trading infrastructure that directly impacts customers.
- Solve challenging reliability and scalability problems in a real-time environment.
- Build world-class observability, automation, and deployment practices.
- Collaborate with experienced engineers in a modern engineering culture.
- Influence reliability strategy and engineering best practices across the platform.
