FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in incident management, operational automation, and production operations within a high-availability environment. Proficient in troubleshooting, communication, and utilizing observability tools to enhance system reliability.
Highest-signal resume keywords
SRE ExperienceIncident Management ToolingScripting Proficiency (Python, Go, Bash)Kubernetes KnowledgeObservability Platform Experience (Datadog)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Troubleshooting SkillsIncident InvestigationOperational AutomationProduction OperationsHigh-Availability Environment ExperienceSLO UnderstandingError BudgetsBurn-Rate AlertingAI/ML-Assisted Operations Interest
Soft Skills
Clear Communication Skills
Tools & Technologies
DatadogJIRAJIRA Service ManagementPagerDutyOpsGenieSlackConfluence
Industry Keywords
Gaming Company ExperienceTechnical OperationsNOC ExperienceDevOps Experience
Tech Stack
Tools & technologiesCloudGoGoogle Cloud PlatformKubernetesPython
About the role
Key responsibilities & impact- Serve as the primary dashboard monitor during your shift
- Triage and investigate production incidents
- Own lower-severity incidents end-to-end from detection through resolution
- Support the TSO Lead during major incidents
- Draft incident communications under TSO Lead direction
- Analyze incident trends, recurring issues, and production bugs
- Compile incident timelines and draft initial PIR documents
- Build and maintain operational automation
- Conduct structured shift handoffs
- Cover for the TSO Lead during vacations, absences, or emergencies
- Publish health reports of critical apps periodically
Requirements
What you’ll need- Previous experience working at a gaming company is required
- 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment
- Strong troubleshooting and investigation skills
- Hands-on experience with Datadog or equivalent observability platform
- Proficiency in at least one scripting language: Python, Go, or Bash
- Clear written and verbal communication skills in English
- Working knowledge of Kubernetes and cloud infrastructure (GCP preferred)
- Understanding of SLOs, error budgets, and burn-rate alerting
- Experience with incident management tooling: JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence
- Experience with or strong interest in AI/ML-assisted operations.
Benefits
Comp & perks- unlimited Flexible Time Off
- Gym membership
- monthly train ticket
- personalized career roadmap for each employee
- professional development through training and educational opportunities
