Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
AH
Aquila Hash Inc.
Data Center Field Engineer
Career Insights for Data Center Technician / Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on New York data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Data Center Technician or Engineer organizes data access, maintenance and access support to a data center. Maybe help to store and organize data, such as an organization's records or financial information. Ensures that users can easily access the information they need and that data are protected from unauthorized access.
$100,189 / year median in New York
+8% projected growth
Job Description
Position Overview The AI GPU Cluster Shift Engineer is responsible for leading on-site operations and maintenance activities for large-scale AI GPU computing clusters during an assigned shift. Each Shift Engineer will lead and coordinate a team of approximately 3-4 Data Center / GPU Operations Technicians, ensuring effective 24×7 operational coverage of GPU servers, network infrastructure and related data center equipment. The Shift Engineer serves as the technical lead and primary operational point of contact during the shift, responsible for monitoring infrastructure health, coordinating incident response, assigning work to technicians, troubleshooting escalated issues, managing vendor support cases and ensuring a complete handover to the next shift. Key Responsibilities Lead and coordinate 3-4 technicians during the assigned shift and manage day-to-day operational activities and work assignments. Monitor the health and status of large-scale AI GPU clusters, including GPU/CPU servers, network switches, optical transceivers and associated infrastructure. Review monitoring alerts, identify infrastructure abnormalities and determine appropriate troubleshooting and escalation actions. Provide technical guidance to technicians for hardware troubleshooting, component replacement, rack-and-stack, cabling, server recovery and other on-site activities. Perform hands-on troubleshooting of server, GPU, network and hardware issues that cannot be resolved by the technician team. Coordinate incident response during the shift and ensure incidents are properly documented, tracked, escalated and communicated to relevant teams. Serve as the primary on-shift escalation point and coordinate with engineering teams, management and other stakeholders as required. Coordinate with OEM and hardware vendors for technical support, hardware replacement and RMA activities, and track cases through resolution. Manage operational changes according to established change-management procedures and Standard Operating Procedures (SOPs). Ensure technicians follow approved SOPs, safety requirements, change procedures and data center operating policies. Verify completion and quality of maintenance activities, hardware replacements, cabling changes and other work performed during the shift. Maintain accurate shift logs, incident records, maintenance records and operational documentation. Conduct a structured shift handover to the incoming Shift Engineer, including active incidents, degraded equipment, pending maintenance, open vendor/RMA cases and other outstanding operational issues. Participate in incident reviews and Root Cause Analysis (RCA), and recommend improvements to SOPs and operational processes. Support spare-parts management and ensure critical replacement components are properly tracked and available for operations. Qualifications Degree or relevant educational background in Computer Science, Information Technology, Engineering, Telecommunications or a related field. 1-2 years of hands-on experience in server, network, data center or infrastructure operations and maintenance. Ability to lead and coordinate a team of 3-4 technicians during an assigned shift, prioritize operational tasks and serve as the primary on-shift escalation point. Hands-on knowledge of x86 server hardware and Linux operating systems. Working knowledge of TCP/IP and data center networking. Hands-on ability to diagnose server hardware failures and replace field-replaceable components. Familiarity with server BMC/OOB management interfaces such as iDRAC, IPMI or equivalent. Ability to interpret monitoring alerts, hardware logs and system information to identify and troubleshoot infrastructure problems. Experience with monitoring platforms such as Zabbix, Prometheus and Grafana is preferred. Experience with GPU servers, RDMA, InfiniBand, RoCE or other high-performance networking technologies is preferred. Good written and verbal communication skills, including the ability to prepare incident reports, RCA documentation, shift reports and technical emails. Ability to communicate effectively with OEM support teams through email, ticketing systems and phone calls. Strong sense of ownership, good organizational skills and the ability to make sound operational decisions during incidents. Willing and able to work 24×7 rotating shifts in an on-site data center environment. Preferred Qualifications Experience operating large-scale AI GPU clusters. Experience supporting NVIDIA or AMD GPU server platforms. Troubleshooting experience with NVIDIA/Mellanox, Broadcom or similar high-performance networking equipment. Experience with Dell, Supermicro or other enterprise GPU/server platforms and OEM support processes. RHCE, CCNA/CCNP or similar industry certifications. Experience with ITIL-based incident, problem and change-management processes. Previous experience coordinating technicians, shift activities or data center operations is a plus.