Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CXM Direct

Application Site Reliability Engineer, SRE

CXM Direct

Application Site Reliability Engineer ensuring reliability of trading systems. Collaborating in a remote team to automate operations and enhance system resilience.

Posted 7/22/2026full-timeRemote • 🇦🇷 ArgentinaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in .NET/C# application reliability, incident response, and operational excellence. Proficient in monitoring, automation, and deployment strategies to enhance system performance and resilience.

Highest-signal resume keywords
Production Support for .NET/C# ApplicationsIncident Response and Root Cause AnalysisPowerShell ScriptingMonitoring with Grafana and PrometheusAWS Infrastructure Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
.NET/C# DevelopmentWindows Server ManagementPython ScriptingBash ScriptingCI/CD Pipeline ManagementTerraformAurora PostgreSQL TroubleshootingService Level Indicators (SLIs)Service Level Objectives (SLOs)Release Automation
Tools & Technologies
GrafanaPrometheusLokiInfrastructure as Code (IaC)Operational Documentation
Industry Keywords
Production OperationsIncident Response ProceduresAlert DesignMetrics and Logging Best PracticesError Budgets

Tech Stack

Tools & technologies
AWSGrafana.NETPostgresPrometheusPythonTerraform

About the role

Key responsibilities & impact
  • Own the day-to-day reliability of .NET/C# services running on Windows.
  • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
  • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
  • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
  • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, AWS infrastructure, CI/CD pipelines and deployments.
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve application operability, resilience, and fault isolation.
  • Automate operational tasks through scripting and infrastructure automation.
  • Create and maintain runbooks, operational documentation, and incident response procedures.
  • Continuously improve monitoring, alert quality, automation, and platform reliability.

Requirements

What you’ll need
  • 3–5 years of experience in .NET/C# applications in production
  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.
  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.
  • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
  • Solid understanding of metrics, logging, tracing, and alerting best practices.
  • Experience with modern CI/CD pipelines.
  • Knowledge of deployment strategies, release automation, and rollback mechanisms.
  • Experience working with AWS.
  • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
  • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), Alert Design, Production Operations.

Benefits

Comp & perks
  • Work on mission-critical trading infrastructure that directly impacts customers.
  • Solve challenging reliability and scalability problems in a real-time environment.
  • Build world-class observability, automation, and deployment practices.
  • Collaborate with experienced engineers in a modern engineering culture.
  • Influence reliability strategy and engineering best practices across the platform.