FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and building AIOps systems, with a strong focus on high-performance distributed systems and ML model-serving infrastructure. Proficient in leveraging diverse database technologies and orchestrating telemetry data for real-time analysis and intelligent alerting.
Highest-signal resume keywords
Expert-Level Proficiency In GoC++ Or Rust ProgrammingKubernetes And Container-Based DeploymentsML Model-Serving Platforms Or MLOps ToolingProduction Distributed Systems Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AIOps System ArchitectureDistributed Systems DesignTelemetry Ingestion And ProcessingData Storage StrategiesPerformance Improvement MetricsModel-Packaging And VersioningReal-Time AnalysisDebugging With Logs And TracesData Movement Across NetworksHigh-Performance Concurrent Architectures
Soft Skills
Problem SimplificationAdaptability In Fast-Moving Environments
Certifications & Qualifications
B.Sc./M.Sc. In Computer ScienceComputer Engineering
Industry Keywords
AI ClustersTelemetry StreamsProduction EnvironmentInternal Tools DevelopmentEnterprise Customer Solutions
Tech Stack
Tools & technologiesC++Distributed SystemsGoKubernetesRust
About the role
Key responsibilities & impact- Architect and build an agentic AIOps system that monitors GPU fleet health, aggregates and correlates telemetry streams, surfaces intelligent alerts, and orchestrates diagnostic workflows and corrective actions
- Research, evaluate, and prototype data storage strategies and data representations across diverse database technologies and modalities
- Design distributed systems for high-density telemetry ingestion, processing, and real-time analysis in large-scale AI clusters
- Instrument services with metrics, logs, and traces for rapid debugging and continuous performance improvement
- Build and own model-serving infrastructure, including packaging, versioning, deployment, and monitoring of AI models in SaaS and on-premises environments
- Contribute to core platform libraries and abstractions that accelerate development across the AIOps engineering team
Requirements
What you’ll need- B.Sc./M.Sc. in Computer Science, Computer Engineering, or a related technical field
- 8+ years of software engineering experience building production distributed systems
- Expert-level proficiency in Go, C++, or Rust, with a focus on high-performance, concurrent architectures
- Solid understanding of Kubernetes and container-based deployments for production services
- Experience deploying, monitoring, and maintaining ML models or data-intensive services in a production environment
- Comfort working in ambiguous, fast-moving environments where the product is still being shaped
- Experience building ML model-serving platforms or MLOps tooling at scale
- Track record of taking systems from prototype to stable, production-grade platform serving real enterprise customers
- Understanding of the full stack, from data movement across the wire to processing in a distributed cluster
- Ability to simplify complex problems and build internal tools or frameworks that empower engineering teams
Benefits
Comp & perks- Competitive salaries
- Generous benefits package
