FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Site Reliability Engineer – Platform
Blue River TechnologyPrincipal SRE scaling Kubernetes infrastructure and Golang services for Blue River’s autonomous robotics platforms. Driving reliability, observability, security, and platform adoption across engineering teams.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and maintaining high-availability infrastructure, with a strong focus on Kubernetes and Golang backend services. Capable of driving architectural decisions, mentoring engineers, and collaborating with cross-functional teams to enhance platform performance and security.
Highest-signal resume keywords
KubernetesGolangCI/CD ToolingCloud SolutionsInfrastructure Architecture
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GolangPythonJavaScriptRustKubernetesTerraformSoftware Design MethodologiesInformation Systems ArchitectureObject-Oriented DesignSoftware Design Patterns
Soft Skills
MentoringCollaborationProblem Resolution
Tools & Technologies
GitHub ActionsArgoCDArgoCD Image UpdaterArtifactoryCloud Vendors
Industry Keywords
High-Availability ApplicationsSaaS Payment ProcessesRisk AssessmentsIntrusion DetectionThreat Feeds
Tech Stack
Tools & technologiesAWSCloudGoJavaScriptKubernetesPythonRustTerraform
About the role
Key responsibilities & impact- Architect, scale, and own essential infrastructure
- Build and maintain a Kubernetes-based platform supporting multiple teams and services
- Build Golang backend services supporting autonomous systems
- Partner with product teams to launch new products on the platform
- Grow high-availability infrastructure while maintaining uptime and other key metrics
- Build tooling for the platform and development teams
- Perform end-to-end performance analysis and implement robust improvements
- Work with cloud vendors and external technical support on upgrades and problem resolution
- Participate in on-call rotation, triage and resolve production incidents, and document root causes and postmortems
- Design and maintain observability infrastructure, dashboards, alerts, and log aggregation
- Collaborate with the security team on regular risk assessments
- Maintain the risk register and implement mitigation plans
- Assess intrusion detection alerts and improve systems that digest threat feeds
- Establish and maintain SaaS payment processes with IT and purchasing teams
- Drive architectural decisions, mentor engineers across teams, and shape the platform's direction
Requirements
What you’ll need- Minimum 8 years of deep experience building and maintaining infrastructure for data-intensive, high-availability applications
- Six years building and maintaining public cloud solutions
- Deep understanding of Kubernetes and Terraform
- Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns
- Deep understanding of securing cloud infrastructure, preferably AWS and Kubernetes
- Deep experience in one or more of Golang, Python, JavaScript, or Rust; Golang preferred
- Deep experience in CI/CD tooling, including GitHub Actions, ArgoCD, ArgoCD Image Updater, and Artifactory
- Visa sponsorship is possible for this position
Benefits
Comp & perks- Annual performance bonus
- Competitive benefit package
- Diversity, Equity, and Inclusion programs
- Recruiting, mentorship, career development, and learning & development programs
- Reasonable accommodation for individuals with disabilities