Job Summary We are seeking a senior GPU Network Engineer to provide Tier 3 operational support for GPU networking environments supporting AI and machine learning workloads. This role will focus on troubleshooting complex issues involving GPU clusters, high-performance networking, RoCE, InfiniBand, and AI infrastructure. The engineer will serve as a technical escalation point and collaborate with network, compute, platform, security, and operations teams to ensure performance, availability, and rapid resolution of production issues. The ideal candidate will have hands-on experience supporting GPU clusters, AI/ML infrastructure, High Performance Computing (HPC) environments, or large-scale low-latency data center networks. Key Responsibilities
- Serve as a Tier 3 escalation resource for GPU network incidents and production issues.
- Troubleshoot complex networking, performance, and connectivity issues impacting GPU clusters.
- Support AI, machine learning, and high-performance computing environments.
- Analyze network performance, congestion, latency, packet loss, and throughput issues.
- Collaborate with security operations, network, compute, platform, and infrastructure teams during incident response.
- Validate and optimize GPU network performance for customer workloads.
- Identify root causes and implement corrective actions for recurring issues.
- Support production turn-up and go-live activities for new GPU environments.
- Participate in operational readiness reviews and knowledge transfer sessions.
- Develop operational runbooks, troubleshooting guides, and support procedures.
- Provide recommendations for capacity planning, scalability, and operational improvements. Required Qualifications
- 5+ years of data center networking experience.
- 2+ years of experience supporting GPU, AI/ML, HPC, or large-scale compute infrastructure environments.
- Experience troubleshooting complex network performance issues in production environments.
- Strong understanding of data center network architecture and operations.
- Experience supporting high-bandwidth, low-latency network fabrics.
- Experience with Arista, NVIDIA, Cisco, or equivalent data center networking technologies.
- Strong understanding of Layer 2 and Layer 3 networking.
- Strong understanding of routing and switching.
- Experience with network monitoring and troubleshooting.
- Experience with incident management and escalation processes.
- Ability to work effectively during critical outages and customer-impacting incidents. Preferred Qualifications
- Experience supporting NVIDIA DGX, HGX, SuperPOD, or similar GPU infrastructure.
- Experience with RDMA, RoCE, InfiniBand, AI network fabrics, or east-west traffic optimization.
- Knowledge of NVIDIA reference architectures.
- Experience with Kubernetes-based AI environments.
- Experience supporting hyperscale, cloud, NeoCloud, or HPC environments.
- Exposure to optical networking, buffer tuning, congestion management, and performance validation.
- Familiarity with network telemetry and performance analytics tools.