Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

E-Solutions Inc.

Senior InfiniBand AI Network :- Santa Clara, CA(Onsite)

Career Insights for Platform Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on California data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Platform Engineer is responsible for the development of platforms that support the needs and use cases of different engineering teams across the organization. Creates reusable tools and workflows to streamline operational needs and facilitate automation tasks, supporting scalability of DevOps practices.

$161,754 / year median in California

Explore Career

Job Description

Senior InfiniBand AI Network 
  • Santa Clara, CA(Onsite) (Santa Clarita, CA, 91350) | 08/26/26 Job Description Role
  • Senior InfiniBand AI Network Location
  • Santa Clara, CA(Onsite) Senior InfiniBand AI Network Position Description
ENGAGEMENT SUMMARY
The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.
WHAT THIS CANDIDATE WILL BE DOING
Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads. Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior. Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance. Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion. Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability. Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source. Define and execute pre-flight and post-change validation workflows for IB cluster readiness. Document recurring fault patterns and create remediation playbooks for operational teams.
WHAT WE NEED TO SEE
7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience. Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics. Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues. Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools. Experience with firmware and driver alignment across HCAs, switches, and Linux hosts. Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.
PREFERRED EXPERIENCE
Experience supporting
DGX, GPU
 superpod, or equivalent AI cluster environments. Familiarity with UFM, telemetry pipelines, and automated fabric validation. Experience correlating IB anomalies with AI training performance outcomes. Senior InfiniBand AI Network 
  • Santa Clara, CA(Onsite)1infini, band United States

Benefits

  • Dental Insurance