Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
JobTailor Logo

See all jobs on JobTailor

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Lambda

HPC Support Engineer

Lambda

Senior HPC Support Engineer at Lambda resolving complex infrastructure and platform issues. Involved in troubleshooting, documentation, and collaboration with engineering teams for permanent fixes.

Posted 7/31/2026full-timeRemote • 🇺🇸 United StatesMid-LevelSenior💰 $122,000 - $162,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in HPC environments, particularly in Linux cluster administration, Kubernetes, and Slurm orchestration. Proficient in troubleshooting complex infrastructure issues, performing root-cause analysis, and utilizing AI tools for operational efficiency.

Highest-signal resume keywords
HPC ExperienceLinux System AdministrationKubernetes Cluster OrchestrationCI/CD ExperienceLog Analysis and Debugging

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
HPC AdministrationLinux SupportCluster AdministrationAI-Assisted ToolsLog AnalysisKernel-Level DebuggingPerformance ProfilingCUDANCCLHigh Throughput Networking
Soft Skills
MentoringCollaborationProblem-SolvingProactive Identification
Tools & Technologies
PrometheusGrafanaDatadogKubernetesSlurm
Industry Keywords
Distributed SystemsAI/ML WorkloadsTCP/IPVPNFirewalls

Tech Stack

Tools & technologies
CloudDistributed SystemsFirewallsGrafanaKubernetesLinuxPrometheusTCP/IP

About the role

Key responsibilities & impact
  • Serve as a senior technical escalation point, troubleshooting the hardest infrastructure and platform issues down to the hardware, driver, or kernel level when needed
  • Quickly and accurately distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
  • Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned
  • Use AI tools effectively to build scripts, automations, or small internal tools that close real operational gaps (no professional development background required)
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
  • Craft clear documentation of solutions and contribute to evolving support procedures
  • Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
  • Take escalations from peers while training and mentoring them in the process
  • Participate in a rotating on-call schedule, owning major incidents and major customer issues
  • Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments

Requirements

What you’ll need
  • 3+ years of hands-on HPC experience in an administration, support, or engineering role.
  • Very strong understanding and experience supporting Linux in a system administration role.
  • Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration.
  • Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast.
  • Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
  • Strong skills in log analysis, debugging kernel-level issues, and performance profiling.
  • Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
  • Experience with high throughput networking technologies(IB/RoCE).
  • Knowledge of distributed AI/ML or HPC workloads.
  • Knowledge of TCP/IP, VPN, and firewalls in cloud environments.
  • Ability to work independently and mentor junior support engineers.

Benefits

Comp & perks
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use