Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Collabera LLC

DevOps Engineer - HPC / EDA / SLURM / Azure

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
78
out of 100
Average of individual scores

Were these scores useful?

Job Description

Job Description

Software Engineer Contract:
Rancho Cordova, California, US Salary Range:

90.00 - 95.00 |

Per Hour Job Code:

372153

End Date:

2026-11-05

Days Left:

29 days, 8 hours left

Job Title:

DevOps Engineer - HPC / EDA / SLURM / Azure (REMOTE)

Duration:

12 Months

100% Remote

Pay Range:

$90-$95/hr

Seeking a Senior DevOps Engineer to provide a 12-month contingent engagement supporting our High Performance Computing (HPC) and Electronic Design Automation (EDA) cloud infrastructure team. This role will work directly within the IT Datacenter (ITDC) organization and is expected to operate independently at a senior level with minimal ramp-up time. The ideal candidate brings strong hands-on experience with Linux HPC environments, infrastructure automation, SLURM workload management, and the Azure cloud environment. What will you do? Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning using Terraform and Ansible. Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers.

Coordinate cross-functionally with

EDA/TD NAND

teams, storage teams, and IDAM to deliver coordinated platform changes.

Administer Azure EDA user environment utilizing Thinlinc (VNC).

Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration.

Ensure multi-version Ansible playbook compatibility across SLES 15.

Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories.

Drive production environment changes through change management workflows using ServiceNow.

Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, and SSSD for Linux/HPC environments.

Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains.

Validate authentication methods and access behavior across HPC compute and storage environments.

Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation.

Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs.

Investigate and remediate operational issues in production Linux services (VNC, AutoFS, Datadog, etc.).

Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence.

What do you need to succeed?

5+ years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role.

Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers.

Demonstrated experience supporting EDA or scientific computing environments.

Strong Ansible & Terraform automation skill with production-grade playbook and role development.

Demonstrated usage and understanding of the Azure cloud compute environment.

Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, Okta).

Experience with NetApp or comparable enterprise storage platforms in HPC contexts.

Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams).

Strong written and verbal communication skill; capable of coordinating across multiple teams.

Preferred Experience:

Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise environment.

Experience migrating configuration artifacts and binaries to Artifactory.

Background in semiconductor, storage, or high-tech manufacturing IT environments The Company offers the following benefits for this position, subject to applicable eligibility requirements: medical insurance, dental insurance, vision insurance, 401(k) retirement plan, life insurance, long-term disability insurance, short-term disability insurance, paid parking/public transportation, paid time off, paid sick and safe time, hours of paid vacation time, weeks of paid parental leave, and paid holidays annually - as applicable. Job Requirement

HPC

EDA

Azure

SLURM

Reach Out to a Recruiter

Recruiter

Email

Phone

PRARTHI MISTRY

prarthi.mistry@collabera.com