Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
EI
E-Solutions Inc.
Senior InfiniBand AI Network :- Santa Clara, CA(Onsite)
Career Insights for Platform Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Platform Engineer is responsible for the development of platforms that support the needs and use cases of different engineering teams across the organization. Creates reusable tools and workflows to streamline operational needs and facilitate automation tasks, supporting scalability of DevOps practices.
$161,754 / year median in California
Job Description
Senior InfiniBand AI Network
- Santa Clara, CA(Onsite) (Santa Clarita, CA, 91350) | 08/26/26 Job Description Role
- Senior InfiniBand AI Network Location
- Santa Clara, CA(Onsite) Senior InfiniBand AI Network Position Description
ENGAGEMENT SUMMARY
The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.WHAT THIS CANDIDATE WILL BE DOING
Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads. Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior. Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance. Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion. Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability. Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source. Define and execute pre-flight and post-change validation workflows for IB cluster readiness. Document recurring fault patterns and create remediation playbooks for operational teams.WHAT WE NEED TO SEE
7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience. Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics. Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues. Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools. Experience with firmware and driver alignment across HCAs, switches, and Linux hosts. Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.PREFERRED EXPERIENCE
Experience supportingDGX, GPU
superpod, or equivalent AI cluster environments. Familiarity with UFM, telemetry pipelines, and automated fabric validation. Experience correlating IB anomalies with AI training performance outcomes. Senior InfiniBand AI Network- Santa Clara, CA(Onsite)1infini, band United States
Benefits
- Dental Insurance