FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive expertise in AWS and Kubernetes, with a strong focus on infrastructure management, disaster recovery, and AI integration. Proven leadership in building and mentoring high-performing teams while ensuring operational excellence and compliance in a high-availability environment.
Highest-signal resume keywords
AWS Architecture DesignKubernetes Cluster ManagementDisaster Recovery PlanningAI Tooling IntegrationInfrastructure-as-Code (Terraform)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AWSKubernetesRDS/Aurora (MySQL and Postgres)Disaster RecoveryInfrastructure-as-CodeCI/CD PracticesIncident ManagementPerformance TuningCost Management (FinOps)Operational Excellence
Soft Skills
LeadershipMentoringCommunicationDecision-MakingCollaboration
Tools & Technologies
TerraformEKSVPC DesignTransit GatewayPrivateLink
Industry Keywords
FintechPCI-DSSSOC 2Audit ReadinessOperational Resilience
Tech Stack
Tools & technologiesAWSCloudKubernetesMySQLPostgresTerraform
About the role
Key responsibilities & impact- Own the infrastructure vision, strategy, and multi-year roadmap: scale today's high-growth fintech platform while continually strengthening its resilience, controls, and audit-readiness.
- Lead, grow, and mentor the infrastructure, platform, and SRE organization, including hiring, career development, on-call health, and building a culture of operational excellence and blameless learning.
- Own reliability end-to-end: define and enforce SLOs and error budgets, mature incident management and postmortem practices, and be accountable for platform availability across the business.
- Manage, and participate in, the on-call rotation, and serve as senior incident commander for high-severity events: leading recovery from major degradations and full outages through rapid, evidence-based triage, decisive action under uncertainty, and clear communication to stakeholders throughout.
- Direct our AWS strategy, including account architecture, IAM and network design, multi-AZ/multi-region posture, service selection, and cost management (FinOps). You will own and defend the cloud budget.
- Own the Kubernetes platform as a product: cluster architecture, upgrade strategy, workload isolation, autoscaling, progressive delivery, and the developer experience of every team that ships on it.
- Own the database tier, centered on Aurora RDS (MySQL and Postgres): availability, performance, capacity, schema and migration safety practices, backup/restore verification, and encryption.
- Design, implement, and continuously test disaster recovery and business continuity: defined RTO/RPO targets per system tier, regular game days and failover exercises, and DR evidence that stands up to auditor scrutiny.
- Champion the AI-boosted SRE transformation: evaluate and deploy AI tooling and agents for incident triage, observability, runbook automation, and toil reduction; set standards for safe, auditable use of AI in production operations; and bring the team along through training and example.
- Partner with Security and Compliance to own infrastructure's role in PCI-DSS and SOC 2: control design and operation, evidence collection, segmentation, vulnerability and patch management, and audit support, with the maturity to meet the expectations of banking partners and financial-industry examinations.
- Drive infrastructure-as-code and platform automation as the default: everything reproducible, reviewed, and recoverable; nothing artisanal.
- Own vendor and technology strategy for the infrastructure domain: build-vs-buy decisions, vendor risk management, contract negotiation, and third-party resilience.
- Communicate crisply with executives, the board, and auditors, translating infrastructure risk, investment, and posture into business terms.
Requirements
What you’ll need- 15+ years of combined experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial depth in infrastructure, including 5+ years leading engineering teams.
- Deep, hands-on expertise with AWS: you have designed and operated production architectures across compute, networking (VPC design, Transit Gateway, PrivateLink), IAM, and multi-account organizations at scale.
- Deep, hands-on expertise with Kubernetes in production: cluster lifecycle management, workload architecture, scaling, and the operational realities of running business-critical services on it (EKS experience strongly preferred).
- Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): high availability, replication, failover, performance tuning, and backup/recovery you have personally verified under pressure.
- Proven ownership of disaster recovery and business continuity for a production platform: you have defined RTO/RPO targets, built the capability to meet them, and run real failover tests, not just written the document.
- Demonstrated AI-forward leadership: you actively use AI tooling in engineering or operations work today, have opinions grounded in practice about where it helps and where it doesn't, and have led (or are visibly leading) a team's adoption of AI-assisted workflows.
- Track record of operating a 24/7, high-availability platform where downtime has direct revenue or customer impact, including mature incident command and postmortem practices.
- Willingness to manage and participate in an on-call rotation, and demonstrated ability to lead recovery from a full production outage: forming and testing hypotheses from logs, metrics, and traces rather than guesswork, making the right call quickly with incomplete information, and knowing when to mitigate first and root-cause later.
- Still technical, by choice: you remain a credible hands-on engineer, comfortable in a terminal, reading dashboards, and reviewing designs, and you expect to stay that way. You will lead the team and work alongside it; this is not a delegation-only role.
- Experience owning significant cloud budgets and driving cost efficiency without sacrificing reliability.
- Strong grounding in infrastructure-as-code (Terraform or equivalent) and modern CI/CD practices.
- Demonstrated ability to hire, develop, and retain strong infrastructure and SRE talent, and to hold a high bar through growth.
- Bachelor's degree in Computer Science or a similar technical field (required).
Benefits
Comp & perks- Unlimited PTO, volunteer hours and sabbatical
- Life, STD/LTD, medical, dental and vision insurance
- Highly discounted LifeTime gym membership
- 401k with match
- Collaborative fun co-workers
- The opportunity to join the fastest growing FinTech alongside a team of motivated and driven individuals
