10+ years in a hands-on infrastructure, HPC , or datacenter engineering role supporting GPU compute at scale. Hands-on experience with
H200/B200/B300
(or comparable) GPU systems: bring up, cabling, firmware/driver management. Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs. Experience with high-performance parallel storage (VAST, DDN, Weka, Lustre, or similar). Kubernetes required;
Run:
ai or similar GPU scheduling/orchestration experience strongly preferred. Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes. Able to lift/move 50+ lbs and perform physical datacenter work (rack/stack/cable/troubleshoot). Eligible to obtain and maintain an active U.S. Top Secret clearance.