Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Technology
Computer Systems Engineer / Architect
Santa Clarita, CA

Find & Apply For Computer Systems Engineer / Architect Jobs in Santa Clarita, California

Browse jobs from a variety of sources below, sorted with the most recently published, nearest to the top. Click the title to view more information and apply online.

Skip to job details
Now viewing: Senior InfiniBand AI Network :- Santa Clara, CA(Onsite)
Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

E-Solutions Inc.

Senior InfiniBand AI Network :- Santa Clara, CA(Onsite)

Job Description

Senior InfiniBand AI Network 
  • Santa Clara, CA(Onsite) (Santa Clarita, CA, 91350) | 08/26/26 Job Description Role
  • Senior InfiniBand AI Network Location
  • Santa Clara, CA(Onsite) Senior InfiniBand AI Network Position Description
ENGAGEMENT SUMMARY
The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.
WHAT THIS CANDIDATE WILL BE DOING
Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads. Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior. Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance. Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion. Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability. Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source. Define and execute pre-flight and post-change validation workflows for IB cluster readiness. Document recurring fault patterns and create remediation playbooks for operational teams.
WHAT WE NEED TO SEE
7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience. Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics. Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues. Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools. Experience with firmware and driver alignment across HCAs, switches, and Linux hosts. Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.
PREFERRED EXPERIENCE
Experience supporting
DGX, GPU
 superpod, or equivalent AI cluster environments. Familiarity with UFM, telemetry pipelines, and automated fabric validation. Experience correlating IB anomalies with AI training performance outcomes. Senior InfiniBand AI Network 
  • Santa Clara, CA(Onsite)1infini, band United States

Benefits

  • Dental Insurance