Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

General Dynamics Information Technology

HPC Infrastructure & Cluster Engineer - TS/SCI with Polygraph

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
75
out of 100
Average of individual scores

Were these scores useful?

Job Description

HPC Infrastructure & Cluster Engineer
  • TS/SCI with Polygraph General Dynamics Information Technology
  • 3.7 Springfield, VA Job Details $119,850
  • $162,150 a year 1 day ago Benefits Internal mobility program Paid holidays 401(k) matching Qualifications Network hardware support Data storage Linux support Automation software Hardware maintenance Hardware support SAN Hardware management Storage management (system administration) Linux administration Full Job Description Clearance Level Top Secret/SCI Category IT Infrastructure and Operations Location Springfield, Virginia ( Onsite Workplace ) Key Skills For Success Cluster Administration Infrastructure Optimization Storage Area Network (SAN) Management REQ#:
RQ227534
Public Trust:
None Requisition Type:
Regular Your Impact Own your opportunity to serve as a critical component of our nation's safety and security. Make an impact by using your expertise to protect our country from threats. Job Description Position Summary We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract at GDIT. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.
Key Responsibilities:
Cluster Administration:
Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
Resource and Job Management:
Configure, maintain, and optimize workload management and orchestration platforms, utilizing the
Run:
AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
Infrastructure Optimization:
Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
Storage and Network Management:
Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
Environment Configuration:
Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
Security and Compliance:
Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Basic Qualifications:
Clearance:
Active TS/SCI with the ability to obtain CI Poly.
Experience:
5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
Location:
Springfield, VA Technical Skills:
Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand). Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g.,
Run:
AI, SLURM
). Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes. Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
Troubleshooting Focus:
Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
Preferred Qualifications:
Familiarity with parallel file systems and high-throughput storage architectures. Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
GDIT IS YOUR PLACE
401K with company match Comprehensive health and wellness packages Internal mobility team dedicated to helping you own your career Professional growth opportunities including paid education and certifications Cutting-edge technology you can learn from Rest and recharge with paid vacation and holidays #RoverGSS Work Requirements Years of Experience 5 + years of related experience may vary based on technical training, certification(s), or degree Certification Travel Required None Citizenship U.S. Citizenship Required Salary and Benefit Information The likely salary range for this position is $119,850
  • $162,150.
This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range. Our Identity Verification Process As part of the hiring process, we will ask you to complete an identity verification process that leverages advanced biometrics and artificial intelligence to ensure authenticity and protect against identity fraud. You are expected to be on camera during virtual interviews. We reserve the right to take your picture to verify your identity and prevent fraud. By proceeding, you authorize the collection, processing, and use of your biometric data for identity verification and security purposes. About Our Work We are GDIT. A global technology and professional services company that delivers technology solutions and mission services to every major agency across the U.S. government, defense and intelligence community. Our 26,000 experts extract the power of technology to create immediate value and deliver solutions at the edge of innovation. We operate across 50+ countries worldwide, offering leading mission-ready capabilities in AI, cloud, cyber and software development. Join our Talent Community to stay up to date on our career opportunities and events at gdit.com/tc. Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans

Benefits

  • Paid Time Off (PTO)
  • 401(k) Plans
  • Dental Insurance