Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Walmart

Senior Site Reliability Engineer

Walmart

Site Reliability Engineer strengthening Walmart’s cloud-powered checkout systems. Building monitoring, automation, incident-response, and resilience tools for millions of daily retail orders.

Posted 8/5/2026full-timeChennai • 🇮🇳 IndiaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering with a strong focus on incident management, performance optimization, and cloud infrastructure. Proficient in developing monitoring solutions and disaster recovery procedures to enhance system reliability and operational efficiency.

Highest-signal resume keywords
Site Reliability EngineeringJavaScript DevelopmentMonitoring and Alerting ToolsDisaster Recovery ProceduresCloud Services Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
JavaNodeJSRESTful ServicesGitMavenJenkinsDockerKubernetesPrometheusGrafana
Soft Skills
Incident TriageRoot Cause AnalysisCollaboration
Tools & Technologies
SplunkAzureGoogle CloudMegaCacheCI/CD
Certifications & Qualifications
SRE Certification
Industry Keywords
Infrastructure ManagementMulti-Cloud EnvironmentsPerformance TuningSLIsSLOs

Tech Stack

Tools & technologies
AngularApacheAzureCloudDockerGrafanaJavaJavaScriptJenkinsKafkaKubernetesLinuxMavenMicroservicesNode.jsPrometheusPythonReactServiceNowSplunkSQLUnix

About the role

Key responsibilities & impact
  • Triage, escalate, and resolve site-impacting production incidents
  • Analyze monitoring graphs, alerts, logs, and system metrics to identify production impacts
  • Design and implement alerting integrations with service API endpoints
  • Monitor and improve adherence to SLIs and SLOs
  • Plan and document disaster recovery procedures for critical applications
  • Optimize Unix/Linux, Java, NodeJS, Tomcat, and Apache performance
  • Develop enterprise monitoring and tooling solutions using Grafana, Splunk, and related technologies
  • Design and develop internal communication workflows and dashboard tools
  • Create and maintain incident-analysis playbooks
  • Handle deployments, post-validations, and back-out procedures
  • Coordinate platform activities including VM upgrades and database maintenance
  • Participate in rotating on-call duties across time zones
  • Perform timely root cause analysis of production issues
  • Develop reusable tooling, libraries, dashboards, and processes to improve customer experience and reduce operational costs
  • Collaborate with developers and platform teams to build observable and resilient systems

Requirements

What you’ll need
  • Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 3 years’ experience in site reliability engineering, site and system administration, infrastructure management, or related area; alternatively, 5 years’ relevant experience without the degree
  • 5+ years of hands-on Site Reliability Engineering, Operations, and Development experience stated in the role description
  • JavaScript, Java, RESTful services, Git, Maven, Jenkins, DevOps, containerization, Docker, Kubernetes, Azure, Google Cloud, Kafka, Azure Cosmos, Azure SQL, MegaCache, CI/CD, Prometheus, Grafana, and Splunk
  • Scripting and software development for automation and self-healing in multi-cloud environments
  • End-to-end understanding of infrastructure, cloud services, platforms, and microservices
  • Ability to triage incidents and distinguish symptoms from causes
  • Knowledge of monitoring and alerting tools, metrics, KPIs, SLIs, SLOs, distributed tracing, and alerting logic
  • Knowledge of disaster recovery procedures, processes, and enterprise disaster recovery systems
  • Knowledge of Unix/Linux, Java, NodeJS, Tomcat, and Apache performance tuning
  • Familiarity with log-centric tooling and time-series data
  • Preferred: master's degree, SRE certification, or 5 years’ relevant experience

Benefits

Comp & perks
  • Incentive awards for performance
  • Maternity and parental leave
  • PTO
  • Health benefits
  • Flexible, hybrid work
  • Training in future skillsets
  • Equal opportunity and inclusive workplace culture