Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Compunnel, Inc.

GPU Network Engineer Operational Support

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
76
out of 100
Average of individual scores

Were these scores useful?

Job Description

Job Summary We are seeking a senior GPU Network Engineer to provide Tier 3 operational support for GPU networking environments supporting AI and machine learning workloads. This role will focus on troubleshooting complex issues involving GPU clusters, high-performance networking, RoCE, InfiniBand, and AI infrastructure. The engineer will serve as a technical escalation point and collaborate with network, compute, platform, security, and operations teams to ensure performance, availability, and rapid resolution of production issues. The ideal candidate will have hands-on experience supporting GPU clusters, AI/ML infrastructure, High Performance Computing (HPC) environments, or large-scale low-latency data center networks. Key Responsibilities
  • Serve as a Tier 3 escalation resource for GPU network incidents and production issues.
  • Troubleshoot complex networking, performance, and connectivity issues impacting GPU clusters.
  • Support AI, machine learning, and high-performance computing environments.
  • Analyze network performance, congestion, latency, packet loss, and throughput issues.
  • Collaborate with security operations, network, compute, platform, and infrastructure teams during incident response.
  • Validate and optimize GPU network performance for customer workloads.
  • Identify root causes and implement corrective actions for recurring issues.
  • Support production turn-up and go-live activities for new GPU environments.
  • Participate in operational readiness reviews and knowledge transfer sessions.
  • Develop operational runbooks, troubleshooting guides, and support procedures.
  • Provide recommendations for capacity planning, scalability, and operational improvements. Required Qualifications
  • 5+ years of data center networking experience.
  • 2+ years of experience supporting GPU, AI/ML, HPC, or large-scale compute infrastructure environments.
  • Experience troubleshooting complex network performance issues in production environments.
  • Strong understanding of data center network architecture and operations.
  • Experience supporting high-bandwidth, low-latency network fabrics.
  • Experience with Arista, NVIDIA, Cisco, or equivalent data center networking technologies.
  • Strong understanding of Layer 2 and Layer 3 networking.
  • Strong understanding of routing and switching.
  • Experience with network monitoring and troubleshooting.
  • Experience with incident management and escalation processes.
  • Ability to work effectively during critical outages and customer-impacting incidents. Preferred Qualifications
  • Experience supporting NVIDIA DGX, HGX, SuperPOD, or similar GPU infrastructure.
  • Experience with RDMA, RoCE, InfiniBand, AI network fabrics, or east-west traffic optimization.
  • Knowledge of NVIDIA reference architectures.
  • Experience with Kubernetes-based AI environments.
  • Experience supporting hyperscale, cloud, NeoCloud, or HPC environments.
  • Exposure to optical networking, buffer tuning, congestion management, and performance validation.
  • Familiarity with network telemetry and performance analytics tools.