Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Amaara Networks

AI Infrastructure Engineer

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
100
out of 100
Average of individual scores

Were these scores useful?

Job Description

AI Infrastructure Engineer Amaara Networks Livermore, CA Job Details Full-time $1 a year 11 hours ago Benefits Paid holidays 401(k) Paid time off Flextime Qualifications Data center experience Data storage Content creation for technical audiences Defect resolution root cause analysis Linux support Technical documentation Defect analysis Hardware maintenance Configuration management Equipment troubleshooting Incident Escalation Computer hardware Incident management operations support Hardware management Storage management (system administration) Problem management Commissioning phase involvement Server management automation Hardware diagnostics Escalation handling Field commissioning Linux administration Technical writing for network engineers Failure analysis Production troubleshooting
Full Job Description AI Infrastructure Engineer Company:
Amaara Networks Location:
Livermore, CA — Hybrid, with regular onsite work at our data center facility
Job Type:
Full-Time Experience Level:
4-7 years
Travel:
Occasional travel, up to 10-15%
Reports To:
Manager, Data Center Operations About Amaara Networks Amaara Networks builds and operates the physical and systems infrastructure that supports modern AI workloads. Our teams deploy, configure, maintain, and troubleshoot high-performance GPU clusters and the networking infrastructure that enables customers to train and run AI models. We are looking for an AI Infrastructure Engineer who enjoys working across hardware, Linux systems, high-performance networking, and GPU infrastructure. This is a hands-on engineering position with responsibility for complex Tier 2/Tier 3 escalations, deployments, troubleshooting, automation, and continuous improvement. Position Summary As an AI Infrastructure Engineer, you will support the full lifecycle of high-performance GPU infrastructure—from physical installation and commissioning through Linux configuration, GPU validation, networking, troubleshooting, and production operations. You will be responsible for resolving complex infrastructure issues, identifying root causes, working directly with vendors and customers, and developing the documentation and automation that makes our infrastructure more reliable and scalable. This role is ideal for someone who enjoys both working in the data center and solving problems from the terminal . Key ResponsibilitiesInfrastructure Troubleshooting & Escalation Own Tier 2 and Tier 3 infrastructure escalations from investigation through resolution. Perform root-cause analysis on server and infrastructure failures using
BMC/IPMI
logs, kernel messages, MCE/EDAC records, thermal and power telemetry, and vendor diagnostics. Identify recurring hardware and infrastructure issues and develop permanent corrective actions. Participate in post-incident reviews and translate findings into monitoring improvements, runbooks, and operational standards. Work directly with hardware vendors on diagnostics, RMAs, firmware issues, and field-service activities. Communicate clearly with customers during incidents and scheduled maintenance windows. Coordinate remote hands and field technicians during work performed at remote sites. Participate in an infrastructure on-call rotation. GPU Infrastructure Deploy, administer, and troubleshoot high-performance GPU compute systems. Support
NVIDIA H100/H200/B200/B300
SXM and PCIe platforms, AMD Instinct, or comparable accelerator platforms. Use tools such as nvidia-smi, DCGM, and DCGM diagnostics to troubleshoot GPU health and performance issues. Analyze XID errors, ECC events, NVLink/NVSwitch health, thermal conditions, and power-related throttling. Install and maintain GPU drivers, CUDA toolkits, Fabric Manager, and container runtimes. Manage software and driver compatibility across the GPU fleet. Perform GPU stress testing, burn-in, NCCL testing, bandwidth testing, and all-reduce benchmarking. Investigate performance issues including straggler nodes, reduced collective bandwidth, PCIe degradation, NUMA configuration, and topology problems. High-Performance Networking Configure and support InfiniBand and RoCEv2 high-performance AI networking environments. Configure switches, manage firmware, and validate network topology. Support InfiniBand subnet management, partitioning, routing, and topology validation. Use tools such as UFM, OFED, ibdiagnet, ibnetdiscover, and perfquery to troubleshoot network issues. Diagnose link flaps, symbol errors, degraded optics, cabling issues, and port-level failures. Configure and troubleshoot RoCE environments, including PFC, ECN, lossless queues, and congestion control. Administer
RDMA/OFED
stacks, NIC/HCA firmware, bonding, VLANs, and management networks. Support enterprise networking, including L2/L3 networking, VLANs, LACP, MLAG/VLT, firewalls, NAT, VPN, DHCP, and DNS. Validate network performance following deployments and infrastructure changes. Maintain accurate network topology, port maps, and infrastructure documentation. Linux Systems Administration Administer Linux systems including RHEL, Rocky Linux, and Ubuntu . Manage provisioning, patching, kernels, drivers, storage, filesystems, networking, users, and access. Operate bare-metal provisioning platforms such as MAAS, Foreman, Ironic, or similar technologies . Troubleshoot PXE, commissioning, enrollment, deployment, and image-related issues. Configure cloud-init, netplan, bonded networking, and storage layouts. Tune Linux hosts for GPU and HPC workloads, including
NUMA, CPU
pinning, IRQ affinity, and I/O configuration. Monitor hosts, GPUs, networking, and environmental conditions. Develop monitoring, dashboards, and alerts to identify issues before they affect customers. Write Python and Bash scripts to automate diagnostics, reporting, and repetitive operational tasks. Data Center Deployment & Physical Infrastructure Rack, cable, configure, and commission high-density GPU servers, storage, and networking equipment. Lead the technical execution of data center deployment windows. Install and terminate copper, fiber, DAC, and AOC cabling for high-speed 100G-800G connections. Plan and verify rack-level power distribution, A/B redundancy, circuit sizing, and load balancing. Contribute to rack and cabinet layouts, including U-space, airflow, weight, cable management, and power requirements. Maintain accurate DCIM and asset records, including rack elevations, cable runs, serial numbers, port maps, firmware inventories, and lifecycle status. Develop and maintain installation standards, validation checklists, and operational runbooks. Qualifications 4-7 years of experience combining data center operations, infrastructure engineering, and Linux systems administration. Strong production Linux administration experience with RHEL, Rocky Linux, or Ubuntu. Experience managing Linux kernels, drivers, storage, networking, and shell-level diagnostics. Hands-on experience operating a bare-metal provisioning platform. Experience with PXE, node commissioning, deployment, and provisioning troubleshooting. Experience using Redfish and IPMI for hardware management, power control, boot configuration, and firmware operations. Demonstrated Tier 2/Tier 3 escalation and incident-resolution experience. Strong server hardware troubleshooting and root-cause analysis skills. Experience with
BMC/IPMI
logs, system event logs, and Linux kernel diagnostics. Experience with high-performance networking, including InfiniBand, RoCE, or high-speed Ethernet . Working knowledge of production L2/L3 networking. Experience supporting GPU or accelerated-computing systems. Understanding of data center fundamentals including power, cooling, airflow, cabling, and rack layouts. Scripting experience with Python or Bash . Familiarity with configuration management tools such as Ansible . Strong technical writing and documentation skills. Ability to communicate directly with customers during infrastructure incidents. Ability to coordinate remote technicians and field-service personnel. Willingness to participate in an on-call rotation and occasional after-hours maintenance windows. Preferred Qualifications Experience supporting large-scale, multi-rack GPU clusters. Experience with
NVIDIA SXM
systems, NVLink/NVSwitch, or rail-optimized network architectures. Deep InfiniBand experience, including UFM, subnet management, ibdiagnet, adaptive routing, or SHARP. Experience with
Slurm, Kubernetes GPU Operators, Run:
ai, or similar workload schedulers. Experience with high-performance storage such as Lustre, GPFS, WEKA, VAST, or NFS over RDMA. Experience with liquid-cooled or high-density data center infrastructure. Experience with Prometheus, Grafana, DCGM Exporter, Elasticsearch, or similar observability platforms.
RHCSA, RHCE, NVIDIA
networking/InfiniBand certifications, CompTIA Server+/Network+, or Uptime Institute certifications. Experience working in a growth-stage environment where processes and standards are being developed. Working Conditions & Physical Requirements This position includes regular work inside an operating data center. Candidates must be able to, with or without reasonable accommodation: Lift and maneuver equipment weighing up to 50 lbs. unassisted and heavier equipment with team or lift assistance. Work on ladders and in confined areas, including under raised floors and in overhead cable trays. Stand, walk, kneel, and reach for extended periods. Work in environments with sustained equipment noise, cold aisles, and limited natural light. Wear required personal protective equipment (PPE). Follow all electrical, data center, and site safety procedures.
Compensation & Benefits Compensation:
Negotiable, depending on experience 401(k) Healthcare reimbursement Paid holidays Flexible time off / flex-time Why Join Amaara Networks? This is an opportunity to work directly with the infrastructure powering modern AI workloads. You'll have the opportunity to work across GPU systems, Linux, high-performance networking, automation, hardware, and data center operations while helping build reliable infrastructure at scale. If you enjoy solving difficult infrastructure problems, working hands-on with cutting-edge technology, and turning recurring problems into permanent solutions, we'd like to hear from you. Apply today to join Amaara Networks as an AI Infrastructure Engineer.
Pay:
$1.00 per year
Benefits:
401(k) Paid time off
Work Location:
In person

Benefits

  • Paid Time Off (PTO)
  • 401(k) Plans
  • Dental Insurance