Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Software Mind

Senior Site Reliability Engineer, Production Support

Software Mind

Senior Site Reliability Engineer supporting the deployment and operations of cloud-native applications on Kubernetes. Focused on maintaining reliability and resolving production incidents for a leading enterprise software company.

Posted 8/1/2026full-timeRemote • 🇨🇦 CanadaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering and DevOps practices, with a strong focus on Kubernetes, cloud-native applications, and production incident management. Proficient in log analysis using Splunk and troubleshooting in Linux environments.

Highest-signal resume keywords
KubernetesSite Reliability EngineeringSplunkCloud-Native ApplicationsLinux

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringDevOpsProduction OperationsLog AnalysisTroubleshootingRoot Cause AnalysisMonitoringCI/CD PipelinesNetworking FundamentalsDistributed Systems
Soft Skills
Excellent Written CommunicationExcellent Spoken Communication
Tools & Technologies
KubernetesSplunkLinux
Industry Keywords
Cloud-Native DeploymentProduction SupportIncident ResponseOperational EfficiencyAutomation

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetesLinuxSplunk

About the role

Key responsibilities & impact
  • Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes.
  • Monitor application health, availability, and performance.
  • Investigate and resolve production incidents using logs, monitoring, and debugging tools.
  • Perform log analysis using Splunk to identify root causes and troubleshoot service issues.
  • Collaborate with software engineers to improve service reliability and operational efficiency.
  • Participate in incident response and production support activities.
  • Assist with first-level debugging of UI-related issues involving Web Components.
  • Contribute to continuous improvements in automation, monitoring, and operational processes.
  • Support CI/CD pipelines and cloud-native deployment practices.

Requirements

What you’ll need
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.
  • Strong hands-on experience with Kubernetes in production environments.
  • Experience supporting cloud-native applications.
  • Experience monitoring production systems and troubleshooting complex incidents.
  • Strong knowledge of Splunk for log analysis and debugging.
  • Experience working in Linux environments.
  • Understanding of networking fundamentals and distributed systems.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent written and spoken English (B2+).

Benefits

Comp & perks
  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives