FREE ACCESS
5,000–10,000 jobs/day
See all jobs on JobTailor
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff HPC Infrastructure Engineer
Guardant HealthStaff HPC Infrastructure Engineer scaling genomics data storage, compute clusters, and networking at precision oncology company Guardant Health. Supporting on-premise and cloud infrastructure.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing HPC clusters, integrating cloud solutions, and optimizing system performance. Proficient in Linux/Unix administration, high-performance networking, and automation tools, with a strong focus on documentation and mentoring.
Highest-signal resume keywords
HPC Cluster ManagementLinux/Unix AdministrationHigh-Performance NetworkingAutomation Tools (Ansible)Cloud Infrastructure (AWS, GCP, Azure)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
HPC Solutions DevelopmentSystem AdministrationTCP/IP NetworkingData Storage ManagementPerformance TuningShell ScriptingPython ProgrammingTechnical DocumentationTroubleshootingCloud Bursting
Soft Skills
MentoringCollaborationCuriositySound Judgment
Tools & Technologies
XDMoDInfiniBandRoCERDMADockerKubernetesSlurmWarewulfRed Hat LinuxGPFS
Certifications & Qualifications
Cisco Certified Network Professional (CCNP)
Industry Keywords
High-Performance Computing (HPC)Cloud InfrastructureCompliance (HIPAA, SOX)Infrastructure EngineeringAutomation Processes
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudDockerGoogle Cloud PlatformKubernetesLinuxPythonTCP/IPUnix
About the role
Key responsibilities & impact- Manage multiple HPC clusters and cluster file systems
- Integrate cloud bursting into HPC abstraction work
- Research, develop, and implement next-generation HPC solutions
- Troubleshoot the production system stack down to source-code level, including shell scripts and Python
- Maintain, monitor, and support infrastructure environments and facilities
- Improve production monitoring capabilities
- Support system reliability and performance improvements
- Support medium- to high-complexity systems and applications with concurrent users, ensuring control, integrity, and accessibility
- Support systems at remote locations, including internationally
- Mentor junior engineers on HPC best practices
- Collaborate with offsite consultants and vendors to maintain, troubleshoot, upgrade, and repair systems
- Represent HPC networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP
- Set up and own XDMoD instances for HPC metrics and monitoring
- Participate in a 24/7 on-call rotation
- Act as technical peer for HPC networking and interconnect initiatives
- Collaborate on HPC Ethernet design, performance tuning, and troubleshooting
- Integrate HPC systems with enterprise bandwidth-on-demand connectivity across sites and cloud infrastructure
- Manage and optimize connectivity to and from HPC systems and global locations
- Partner with the storage engineer and MSP on HPC storage architecture and integration
- Serve as technical point of contact for the MSP storage relationship, helping define SLAs, validate delivery, and escalate technical issues
- Support transition of day-to-day storage operations to the MSP while maintaining performance and reliability
Requirements
What you’ll need- Bachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experience
- Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking
- Hands-on experience with automation tools, such as Ansible or equivalent technologies
- Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments
- Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments
- Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments
- Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets
- Strong experience creating and maintaining system administration and technical documentation
- Ability to participate in a 24/7 on-call rotation
- Background screening including criminal history required
- Preferred: Cisco Certified Network Professional (CCNP) certification
- Preferred: Experience with Arista networking, GPFS, Slurm, Warewulf, Red Hat Linux, cloud bursting, wide area file systems, Docker, Apptainer, Kubernetes, and HIPAA/SOX-compliant infrastructure
- Demonstrated curiosity, sound judgment, and ability to critically evaluate and responsibly leverage AI-enabled tools according to company policies, ethical standards, and regulatory requirements
Benefits
Comp & perks- Hybrid work model with defined in-person/onsite collaboration days and work-from-home days for individual-focused time
- Flexible work-life balance
- Reasonable accommodations during hiring processes
- Confidential handling of applicant information
- Equal Opportunity Employer