FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer – Application Reliability, Kubernetes, GCP, SQL
CiscoSenior engineer ensuring reliability of Cisco’s enterprise AI applications through observability, automation, and self-healing. Building Kubernetes, GCP, BigQuery, LangGraph, and Python solutions.
Posted 9/3/2026full-timeSan Jose • California, North Carolina • 🇺🇸 United StatesSenior💰 $203,000 - $258,600 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in application reliability and observability, leveraging strong Python development skills and deep SQL knowledge with BigQuery. Capable of implementing SLI/SLO frameworks and driving a reliability culture across engineering teams.
Highest-signal resume keywords
Python DevelopmentBigQuery SQL ExpertiseGCP ExperienceApplication-Level SLI/SLO FrameworksAgent Evaluation Harnesses
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Application ReliabilityObservability SystemsAnomaly DetectionRoot Cause AnalysisFeature FlagsAutomated RollbackComplex SQL QueriesSchema DesignDistributed TracingDebugging Skills
Soft Skills
CollaborationLeadershipCommunication
Tools & Technologies
LookerBigQueryBigTableGKECloud LoggingCloud TraceCloud MonitoringLangGraph
Certifications & Qualifications
Bachelor's Degree in Computer ScienceMaster's Degree in Engineering
Industry Keywords
AI-Powered ApplicationsProduction OperationsReliability EngineeringAIOpsEvent-Driven Systems
Tech Stack
Tools & technologiesBigQueryCloudGoogle Cloud PlatformKubernetesPythonSQL
About the role
Key responsibilities & impact- Own the reliability of AI-powered applications and features from the user’s perspective
- Define, implement, and enforce feature-level SLIs, SLOs, and error budgets for APIs, RAG systems, AI agents, and user-facing applications
- Build and maintain application observability systems using Looker dashboards on BigQuery and BigTable
- Provide visibility into feature health, error patterns, and usage trends for developers, product managers, and leadership
- Design and build LangGraph-based agents for automated issue identification and remediation
- Implement anomaly detection, root cause diagnosis, auto-rollback, feature-flag kill switches, and self-healing workflows
- Develop agent evaluation harnesses for benchmarking, multi-step workflow testing, non-deterministic outputs, and regression testing
- Write complex BigQuery SQL for usage trend analysis, anomaly detection, and operational analytics
- Design BigQuery table schemas optimized for observability and debugging
- Analyze application usage trends and adoption metrics to identify reliability risks, capacity needs, and degraded user experiences
- Partner with application development teams to embed deployment safety, structured logging, and distributed tracing practices
- Lead application-level incident response, root cause analysis, and blameless postmortems
- Build Python tooling and automation to reduce mean time to detect and resolve application-layer issues
- Apply emerging AI techniques to improve platform reliability and developer productivity
- Collaborate with application developers, data engineers, infrastructure SREs, security, compliance, product teams, and leadership
Requirements
What you’ll need- 10+ years of experience in software engineering with significant focus on reliability, observability, or production operations
- Bachelor's or Master's Degree in Computer Science, Engineering, or a related technical discipline
- Strong Python development skills, including production tooling, automation, and agent-based systems
- Production GCP experience deploying and managing applications on GKE (Kubernetes)
- Deep SQL expertise with BigQuery, including complex queries, window functions, schema design, and cost optimization
- Hands-on experience with BigTable or an equivalent high-throughput operational data system
- Experience designing and operating application-level SLI/SLO frameworks, burn-rate alerting, and error budget policies
- Strong application-layer debugging skills, including distributed tracing, profiling, structured log analysis, and dependency mapping
- Experience building agent evaluation harnesses
- Familiarity with A2A protocols, streaming architectures, and event-driven systems
- Experience with feature flags, canary deployments, progressive rollouts, and automated rollback
- Experience with GCP observability services such as Cloud Logging, Cloud Trace, and Cloud Monitoring
- Exposure to AIOps concepts including ML-driven anomaly detection, automated root cause analysis, and intelligent alerting
- Experience driving reliability culture across engineering teams
- Active engagement with the evolving AI ecosystem
- Hands-on GenAI application development experience with LangGraph, agent engineering, prompt design, and agentic workflows
- Experience building Looker dashboards and LookML models for operational observability
Benefits
Comp & perks- Medical, dental and vision insurance
- 401(k) plan with a Cisco matching contribution
- Paid parental leave
- Short- and long-term disability coverage
- Basic life insurance
- Grants of Cisco restricted stock units, subject to eligibility and continued employment
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- Paid day off for employee’s birthday
- Paid year-end holiday shutdown
- 4 paid days off for personal wellness
- 16 days of paid vacation time per full calendar year for non-exempt employees
- Flexible vacation time off program with no defined limit for eligible exempt employees
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward
- Additional paid time away for critical or emergency family issues
- Optional 10 paid volunteer days per full calendar year
- Annual bonuses for non-sales roles, subject to Cisco’s policies
- Performance-based incentive pay for sales-plan employees