FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in cloud infrastructure management, particularly with AWS, and possesses a strong background in designing reliable, event-driven systems. Proven ability to lead technical direction, mentor engineering teams, and implement robust reliability strategies across platforms.
Highest-signal resume keywords
AWS ExpertiseEvent-Driven System DesignInfrastructure as CodeTechnical LeadershipMonitoring and Alerting Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Event-Driven SystemsAsynchronous CommunicationAWS EC2AWS VPCAWS IAMAWS S3AWS RDSTerraformKubernetesDocker
Soft Skills
Technical MentorshipCommunicationProblem-Solving
Tools & Technologies
KafkaNATSRabbitMQDatadogGremlinChaos MeshAWS FISSQLPostgreSQLMongoDB
Industry Keywords
Reliability EngineeringSLOsSLIsError BudgetsDistributed Systems
Tech Stack
Tools & technologiesAWSCloudDistributed SystemsDockerEC2GoKafkaKubernetesMongoDBNoSQLPostgresPythonRabbitMQRedisSQLTerraform
About the role
Key responsibilities & impact- Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
- Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
- Define SLO culture, error-budget policy, and incident practices across engineering teams
- Design and own durable, reliable asynchronous messaging for inter-service communication
- Own cloud infrastructure and automate provisioning with Infrastructure as Code
- Ensure the platform scales reliably as transaction volume grows
- Build monitoring, tracing, and alerting systems for platform health
- Serve as senior escalation point for difficult production incidents
- Run blameless postmortems and root-cause analyses that produce permanent fixes
- Conduct continuous fault injection and resilience experiments
- Mentor senior and mid-level engineers and raise organization-wide reliability standards
Requirements
What you’ll need- 7+ years of experience
- Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
- Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
- Experience migrating systems from synchronous to asynchronous communication
- Deep AWS experience with EC2, VPC, IAM, S3, and RDS
- Strong networking fundamentals
- Infrastructure as Code experience with Terraform or Pulumi
- Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
- Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
- Track record defining and operating SLOs, SLIs, and error budgets across services
- Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
- Distributed systems debugging experience
- Comfortable coding automation and tooling in Go, Python, or similar
- Solid SQL and PostgreSQL knowledge
- NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
- Proven technical leadership, architecture influence across teams, and engineering mentorship
- Advanced written and spoken English proficiency
Benefits
Comp & perks- Competitive Compensation
- Remote Work — you can work from everywhere
- Home Office Bonus — a one-time allowance to set up your ideal home office
- Work Equipment
- Stock Options
- Health Plan wherever you are
- Flexible Days Off
- Language, Professional, and Personal Growth courses
