Find Jobs Near You – Available Work in Your Location
Technology
Computer Systems Engineer / Architect
Santa Clarita, CA
Find & Apply For Computer Systems Engineer / Architect Jobs in Santa Clarita, California
Browse jobs from a variety of sources below, sorted with the most recently published, nearest to the top. Click the title to view more information and apply online.
Senior InfiniBand AI Network :- Santa Clara, CA(Onsite)
Pay: Not specified
Posted: 2 weeks ago
Location: Santa Clarita, CA (Onsite)
Last Updated: 4 days ago
Hours: Full-Time
Expires: 10/12/2026
Job Description
Senior InfiniBand AI Network
Santa Clara, CA(Onsite) (Santa Clarita, CA, 91350) | 08/26/26 Job Description Role
Senior InfiniBand AI Network Location
Santa Clara, CA(Onsite) Senior InfiniBand AI Network Position Description
ENGAGEMENT SUMMARY
The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.
WHAT THIS CANDIDATE WILL BE DOING
Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads.
Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior.
Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance.
Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion.
Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability.
Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source.
Define and execute pre-flight and post-change validation workflows for IB cluster readiness.
Document recurring fault patterns and create remediation playbooks for operational teams.
WHAT WE NEED TO SEE
7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience.
Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics.
Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues.
Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools.
Experience with firmware and driver alignment across HCAs, switches, and Linux hosts.
Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.
PREFERRED EXPERIENCE
Experience supporting
DGX, GPU
superpod, or equivalent AI cluster environments.
Familiarity with UFM, telemetry pipelines, and automated fabric validation.
Experience correlating IB anomalies with AI training performance outcomes.
Senior InfiniBand AI Network
Santa Clara, CA(Onsite)1infini, band United States