# Senior HPC Engineer - Fleet Engineering at Lambda

- Company: Lambda
- Status: Open
- Workplace: Hybrid
- Location: San Francisco Office (Fremont St) · Remote · San Jose Office (First St) · Remote, USA · Bellevue Office
- Level: Senior
- Discipline: DevOps & SRE
- Employment: Full-time
- Salary: $227K–356K
- Posted: 2026-09-01
- Skills: Data Center Business, Fleet Engineering
- Apply: https://jobs.ashbyhq.com/lambda/d4e6f207-73c3-43d1-aa4b-d840ad2254d4

## Description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.



If you'd like to build the world's best AI cloud, join us.




What You’ll Do

 - Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively

 - Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible

 - Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand

 - Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely

 - Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teams

 - Participate in on-call rotations and lead incident response for cluster-level problems

 - Contribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiency

You

 - 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role

 - Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware per

More Lambda roles: https://deviantjobs.com/companies/lambda
