FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in operating and supporting critical production environments on AWS, with a strong focus on Kubernetes and incident management. Proficient in diagnosing issues across various layers and implementing improvements for performance and cost optimization.
Highest-signal resume keywords
AWS Services (EC2, RDS, S3, SQS, SNS, CloudFront, API Gateway)Kubernetes and Amazon EKS TroubleshootingIncident Management and Root Cause AnalysisDynatrace for Alert Analysis and Event CorrelationAPI Gateway and Service Mesh Solutions
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesAmazon EKSAWS ServicesDynatraceIncident ManagementRoot Cause AnalysisAPI GatewayService MeshCloud MonitoringCost Optimization
Soft Skills
Strong Communication SkillsAutonomyProactivity
Tools & Technologies
DynatraceDatadogSQSSNSAPI GatewayKongKrakenDEnvoyQuickSightOpenMetadata
Industry Keywords
Financial SectorPaymentsOpen FinanceFinOps Practices
Tech Stack
Tools & technologiesAWSEC2Kubernetes
About the role
Key responsibilities & impact- Operate and support critical production environments on AWS
- Monitor, analyze and resolve incidents using Dynatrace, Datadog and synthetic monitoring
- Perform advanced troubleshooting in Kubernetes/EKS, including pods, OOM, CPU and memory consumption, resource limits, HPA and scaling
- Investigate failures involving APIs, networking, databases, messaging and data pipelines
- Diagnose issues with SQS and SNS, including DLQs and stuck/retained messages
- Conduct root cause analyses and implement corrective and preventive actions for recurring incidents
- Support the stability and performance of API Gateways, service mesh and integration components
- Communicate incidents, risks and progress to the client, ensuring SLAs are met
- Propose improvements for availability, observability, capacity and AWS cost optimization
- Work autonomously and perform technical escalation when necessary
Requirements
What you’ll need- Experience supporting critical 24x7 production environments
- Proficiency with Kubernetes and Amazon EKS, including troubleshooting, resource management, scaling and high availability
- Experience with AWS services, especially EC2, RDS, S3, SQS, SNS, CloudFront and API Gateway
- Experience with Dynatrace for alert analysis, event correlation, root cause identification and tuning
- Ability to diagnose problems across different layers: infrastructure, containers, applications, APIs, networking, databases and messaging
- Experience in incident management, root cause analysis and remediation of recurring issues
- Autonomy, proactivity and strong communication skills with technical teams and clients
- Preferred:
- Experience with Kong, KrakenD, Envoy or other API Gateway and service mesh solutions
- Knowledge of Keycloak, AWS IAM and authentication/SSO solutions
- Familiarity with QuickSight, OpenMetadata, Budibase or CKAN
- Knowledge of FinOps practices and AWS cost optimization
- Previous experience in the financial sector, payments or Open Finance
Benefits
Comp & perks- Flexible working hours
- Educational incentives (partnerships with educational institutions)
- Paid vacation
- TotalPass
- Birthday off
- Health insurance
- Dental insurance
- Maternity leave
- Paternity leave
- Reimbursement for AWS certifications
