FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Software Engineer, Infrastructure
AnyscaleSoftware Engineer developing infrastructure tools and services for scalable AI applications at Anyscale. Focusing on control plane and data plane development in a hybrid environment.
Posted 7/21/2026full-timeSan Francisco • California • 🇺🇸 United StatesMid-LevelSenior💰 $215,000 - $275,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing distributed systems for AI/ML workloads, with a strong focus on cloud-native technologies and Kubernetes deployments. Proficient in programming with Go and Python, and experienced in managing observability and resource management for scalable infrastructures.
Highest-signal resume keywords
Cloud-Native TechnologiesKubernetes-Based DeploymentsDistributed Systems DesignGo ProgrammingPython Programming
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed System DevelopmentAI/ML Workload OptimizationControl Plane Component OptimizationResource Management SystemsContainer Image ManagementDependency ResolutionNetworking FundamentalsSecurity MechanismsLinux Kernel KnowledgeFile Systems Understanding
Soft Skills
CollaborationTroubleshootingCode Review Participation
Tools & Technologies
AWSAzureGCPKubernetesPrometheusGrafana
Industry Keywords
AI InfrastructureScalabilityObservabilityHigh-Quality Production CodeOn-Call Support
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonRay
About the role
Key responsibilities & impact- Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
- Optimize control plane components for large-scale, distributed AI/ML workloads
- Build intelligent scheduling and resource management systems for heterogeneous compute clusters
- Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
- Support and optimize accelerator integration (e.g., GPUs, TPUs).
- Handle container image management and dependency resolution for distributed workloads
- Participate in code reviews, design and architecture discussions
- Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
- Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
Requirements
What you’ll need- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
- 3+ years of experience writing high-quality production code
- Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
- Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
- Deep understanding of networking, security, and authentication mechanisms in cloud environment
- Familiarity with observability stacks (Prometheus, Grafana etc)
- Proficiency in Go and Python
- Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Benefits
Comp & perks- Stock Options
- Healthcare plans, with premiums covered by Anyscale at 99%
- 401k Retirement Plan
- Wellness & Education Stipend
- Paid Parental Leave
- Fertility Benefits
- Paid Time Off
- Commute reimbursement
- 100% of in office meals covered