FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Lead Site Reliability Engineer – Cloud
ScalingoLead Site Reliability Engineer ensuring reliability and performance of Scalingo's cloud platform. Leading an SRE team and shaping practices for technical excellence in a growing tech startup.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates strong expertise in cloud environments and distributed infrastructure, with proficiency in observability practices and a solid understanding of containerized environments. Capable of providing technical leadership, mentoring others, and ensuring high service quality while optimizing resource usage and scalability.
Highest-signal resume keywords
Cloud Environments ExpertiseObservability PracticesInfrastructure as CodeProduction Database ManagementTechnical Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Incident Response DesignPerformance AnalysisResource OptimizationContainerized EnvironmentsDatabase ReliabilityAutomationOperational Security AwarenessAI Tools UtilizationComplex Incident DiagnosisService Level Agreements (SLAs)
Soft Skills
Clear CommunicationCross-Functional CollaborationTechnical CuriosityComposureBlameless Mindset
Tools & Technologies
Monitoring ToolsMetrics ToolsLogging ToolsAlerting Tools
Industry Keywords
Production Systems StabilityIncident RetrospectivesUser Impact FocusOperational Challenges
Tech Stack
Tools & technologiesCloud
About the role
Key responsibilities & impact- Ensure the stability, availability and resilience of production systems
- Anticipate failures and design effective incident response processes
- Industrialize and automate platform operations
- Maintain a high level of service quality for our customers and comply with contractual commitments (SLAs)
- Analyze performance, identify bottlenecks and propose improvements to optimize resource usage and scalability
- Define, implement and improve observability tools (monitoring, metrics, logs, alerting) with a proactive approach
- Provide level-3 customer support in coordination with support teams and according to SLAs
- Lead and facilitate incident retrospectives (post-mortems), identify root causes and define sustainable corrective actions
Requirements
What you’ll need- Strong expertise in cloud environments and distributed infrastructure
- Proficiency in observability practices (logs, metrics, alerting) and a structured approach to diagnosing complex incidents
- Solid understanding of containerized environments and their operational challenges
- Proven skills with production databases: reliability, backups, restores, replication and scaling
- Hands-on experience with Infrastructure as Code and environment automation
- Awareness of operational security concerns
- Comfortable using AI tools to improve day-to-day efficiency
- Ability to operate in complex, changing or uncertain contexts with rigor and reliability
- Comfortable prioritizing work, including during incident situations
- Clear and structured communication, a preference for cross-functional collaboration and knowledge sharing
- A blameless mindset, technical curiosity, composure and a focus on user impact
- Ability to provide technical leadership, mentor others and advance collective practices
Benefits
Comp & perks- Fully remote with one trip per quarter (Strasbourg or another city)
- Company events: one annual offsite and regular afterworks/social gatherings
- Telework allowance (€57.60)
- Meal vouchers (Ticket Restaurant) (€11.52 per voucher) and Swile card with its benefits
- Flexible hours under a fixed-hours agreement (including RTT)
- Linux laptop provided
- Budget contribution for additional equipment