FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Distinguished Software Engineer
WalmartDistinguished Software Engineer providing technical leadership and architectural design for Walmart's ML data and feature store platform. Delivering scalable and reliable data systems for diverse applications.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and optimizing ML data and feature store platforms, with a strong focus on data quality, observability, and reliability. Proficient in Java and Python SDK design, distributed systems, and high-performance data processing pipelines.
Highest-signal resume keywords
Java SDK/Library DesignPython SDK Design for Data ScienceDistributed Systems FundamentalsObservability Tooling ExperienceKubernetes Deployment at Scale
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
JavaPythonData ModellingBatch Data PipelinesLow-Latency Data StoresGRPC Service DesignCI/CD Pipeline DesignObservability FrameworksArchitectural StandardsReliability Engineering
Soft Skills
MentorshipCross-Functional CollaborationTechnical Authority
Tools & Technologies
SparkHiveAirflowOpenTelemetryPrometheusGrafanaKubernetesGCPAzureRedis
Industry Keywords
Machine LearningData EngineeringFeature StoreDistributed SystemsObservabilityData Quality StandardsMulti-Cloud StrategiesProduction Software DevelopmentLatency OptimizationSchema Design
Tech Stack
Tools & technologiesAirflowAzureCassandraCloudDistributed SystemsGoogle Cloud PlatformGrafanaGRPCJavaKubernetesPrometheusPythonRedisSpark
About the role
Key responsibilities & impact- Define and own the short term and long-term technical vision for the ML data and feature store platform, spanning offline pipelines, online serving, and the infrastructure that connects them
- Identify and systematically retire architectural debt across the platform — prioritising by production impact, latency, reliability, and engineering velocity
- Serve as the primary technical authority in cross-functional design reviews involving ML engineering, data engineering, platform, and product teams
- Influence the broader engineering organisation through published architectural standards, internal RFCs, and mentorship of senior and staff engineers
- Architect end-to-end feature store data systems — from data ingestion and feature computation through materialization into online stores and delivery to model inference
- Define and Design high performant, scalable, reliable Data and Control Mgmt planes for both online and offline data processing pipelines
- Define and govern data quality standards, SLOs, and observability frameworks across the platform
- Drive platform decisions around schema design, storage backend selection, consistency trade-offs, and multi-cloud data replication strategies
- Design and optimise large-scale batch data pipelines using Spark, Hive, or equivalent
- Establish pipeline reliability and observability standards: data freshness SLOs, quality checks at ingestion, deduplication, late-arrival handling, and lineage tracking
- Provide guidance on orchestration and scheduling of batch workloads (Airflow or equivalent)
- Architect and optimise online feature serving infrastructure targeting high QPM at consistently low latency
- Define online serving reliability patterns and establish capacity planning and load testing practices for online stores
- Lead the design and evolution of the platform's Java and Python SDKs — from the current thick SDK to a thin SDK
Requirements
What you’ll need- Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 6 years’ experience in software engineering or related area
- 8 years’ experience in software engineering or related area
- 10+ years of production software development experience, with demonstrable ownership of complex, distributed systems
- Java (Expert): Concurrency primitives (AtomicReference, CAS, CompletableFuture, ForkJoin), SDK/library design, API versioning, backward compatibility guarantees, and dependency management
- Python (Strong): Production-grade library and SDK design for data science or ML engineering consumers; idiomatic async client patterns; packaging and dependency management
- Strong data modelling instincts across relational, wide-column, key-value, and columnar storage paradigms
- Experience with observability tooling: distributed tracing (OpenTelemetry), structured logging, metrics (Prometheus/Grafana or equivalent), and SLO-based alerting
- Experience deploying and operating production services on Kubernetes at scale (GKE, AKS, or equivalent)
- Familiarity with GCP and/or Azure infrastructure: managed storage services, multi-cloud networking, IAM/RBAC, and cloud-native observability tooling
- Understanding of CI/CD pipeline design for data and ML workloads — including pipeline testing strategies, staged rollouts, and data migration patterns
- Proven experience designing and operating services at high QPM with p99 latency targets in the low-millisecond range, across distributed, multi-region infrastructure
- Deep expertise in low-latency data stores: Cassandra and Redis
- Expert-level understanding of distributed systems fundamentals: CAP theorem, consistency models, replication topologies, partitioning strategies, and eventual consistency trade-offs
- Deep expertise in reliability engineering for online systems: circuit breakers, exponential backoff with jitter, bulkhead patterns, rate limiting, local vs. remote DC failover, and graceful degradation under store-side pressure
- Experience with gRPC service design at scale: service contract definition, proto schema evolution, deadline propagation, streaming patterns, and observability
Benefits
Comp & perks- incentive awards for your performance
- maternity and parental leave
- PTO
- health benefits
- flexible work arrangements