Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Ambient Security

Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Ambient Security

Software Engineer building inference, evaluation, and serving infrastructure for Ambient.ai’s physical security AI platform. Optimizing computer vision, LLM, and multimodal systems for real-time enterprise use.

Posted 9/11/2026full-timeRedwood City • California • 🇺🇸 United StatesJuniorMid-Level💰 $168,000 - $205,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and optimizing AI infrastructure for real-time computer vision and multimodal inference workloads, with a strong focus on model evaluation, deployment, and performance optimization. Proficient in collaborating with cross-functional teams to enhance model-serving architectures and ensure production readiness.

Highest-signal resume keywords
Python ProgrammingMachine Learning InfrastructureInference OptimizationModel-Serving FrameworksCloud Infrastructure

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI Infrastructure DesignDeep Learning ModelsScalable Systems DevelopmentEvaluation FrameworksData Engine DesignBatching and CachingQuantization TechniquesGPU UtilizationModel Quality MeasurementProduction AI Systems
Soft Skills
CollaborationCommunicationProblem-SolvingOwnership MindsetAdaptability
Tools & Technologies
VLLMTriton Inference ServerCUDAPyTorchTensorRTONNXContainersOrchestrationDistributed SystemsVideo Understanding
Industry Keywords
Computer VisionLLMsLVMsMultimodal AIRAG PipelinesEmbedding ModelsModel CompressionSpeculative DecodingPruningVector Databases

Tech Stack

Tools & technologies
CloudDistributed SystemsPythonPyTorch

About the role

Key responsibilities & impact
  • Design, build, and maintain AI infrastructure for real-time computer vision, LLM, LVM, and multimodal inference workloads
  • Build scalable systems for state-of-the-art models across large volumes of video and sensor data
  • Optimize inference for latency, throughput, GPU utilization, reliability, and cost
  • Develop evaluation harnesses and benchmarking systems for model quality, system performance, regressions, and production readiness
  • Build infrastructure for continuous model evaluation, experimentation, and deployment
  • Partner with research scientists to productionize advances in computer vision, LLMs, LVMs, RAG, and multimodal AI
  • Improve model-serving architecture through batching, caching, routing, quantization, model parallelism, and hardware utilization
  • Develop data engines and feedback loops for training data collection, model behavior evaluation, and continuous AI improvement
  • Create observability, monitoring, and debugging tools for production AI systems
  • Define best practices for deploying, evaluating, and operating AI systems in enterprise environments
  • Collaborate with research scientists, product engineering, infrastructure teams, and stakeholders

Requirements

What you’ll need
  • 2+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems
  • BS/MS in Computer Science or a related technical field, or equivalent practical experience
  • Strong programming background, especially in Python, with solid software engineering fundamentals
  • Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment
  • Hands-on experience running deep learning models in production, ideally including LLMs, LVMs, vision-language models, or multimodal models
  • Strong understanding of inference optimization, including batching, caching, quantization, parallelism, memory optimization, GPU utilization, and latency reduction
  • Experience with model-serving frameworks such as vLLM, Triton Inference Server, or similar technologies
  • Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model-quality measurement systems
  • Strong background in machine learning and deep learning
  • Experience designing data engines or pipelines for training and evaluation data
  • Familiarity with LLMs, LVMs, RAG pipelines, embedding models, or multimodal models in production applications
  • Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU-based workloads
  • Nice to have: large-scale GPU infrastructure, CUDA, NCCL, PyTorch, TensorRT, ONNX, video understanding, model compression, speculative decoding, distillation, pruning, prompt evaluation, vector databases, re-rankers, search infrastructure, or internal ML platforms
  • Strong collaboration and communication skills
  • Proactive problem-solving ability, ownership mindset, and adaptability

Benefits

Comp & perks
  • Stock options for regular full-time employees
  • Comprehensive health and welfare package: Medical, Dental, Vision, Life, EAP, Legal Services, 401k plan
  • Flexible time off
  • Winter Break time off between Christmas and New Year's for most roles, depending on customer demand
  • Latest tech and swag delivered to your door
  • Opportunities to connect with co-workers
  • Hiking, food, and music activities