Skip to main content We use cookies to provide you with the best experience on our website, to improve usability and performance and thereby improve what we offer to you. Our website may also use third-party cookies to display advertising that is more relevant to you. If you want to know more about how we use cookies, please see our "Required cookies allow us to offer you the best possible experience when accessing and navigating through our website and using its features.
Some examples include:
session cookies needed to transmit the website, authentication cookies, and security cookies. The website will not function properly without these cookies." Please see our Privacy Policy Read Full Privacy Message
Decline
Accept Cookies
Sign In
Home
Search for Jobs
Join Our Talent Community!
Business Title Sr IT Infrastructure Specialist page is loaded
Business Title Sr IT Infrastructure Specialist locations
Heathrow, FL - USA (Office)
Roseville, MN - USA (Office)
time type
Full time
posted on
Posted Yesterday
job requisition id
R04733 Cohesity is a leader in AI-powered data security and management. Aided by an extensive ecosystem of partners, Cohesity makes it easy to secure, protect, manage, and get value from data — across the data center, edge, and cloud. Cohesity helps organizations defend against cybersecurity threats with comprehensive data security and management capabilities, including immutable backup snapshots, AI-based threat detection, monitoring for malicious behavior, and rapid recovery at scale. We've been named a Leader by multiple analyst firms and have been globally recognized for Innovation, Product Strength, and Simplicity in Design. Join us on our mission to shape the future of our industry. Senior IT Infrastructure Engineer About the Role We are seeking a Senior IT Infrastructure engineer who thrives in dynamic environments and can rapidly learn new platforms, embrace AI-assisted operations, and contribute across both lab and production infrastructure. Success in this role requires strong technical depth, operational ownership, collaboration, and a bias toward automation and continuous improvement. You'll work across Linux systems, Kubernetes clusters, kubevirt-based virtualization, and enterprise SAN storage. You will be expected to use AI tools to move faster, reduce toil, and catch issues earlier than a purely manual workflow would allow. This is not an "AI-optional" role. We expect modern infrastructure engineers to pair deep systems knowledge with AI-assisted workflows for scripting, troubleshooting, documentation, and operational analysis. The judgment and systems expertise are still yours; AI is a force multiplier, not a replacement for understanding what's happening under the hood. What You'll Do Design, deploy, and maintain Linux server infrastructure (RHEL/Rocky/Ubuntu and similar) across production and non-production environments
Operate and scale Kubernetes clusters: deployments, networking, storage integration, upgrades, and troubleshooting
Manage kubevirt-based virtualization environments: VM lifecycle, resource allocation, live migration, performance tuning, and host maintenance
Administer and optimize SAN storage systems (Pure Storage, Dell, HPE, and/or NetApp), including provisioning, performance, snapshots, and capacity planning
Contribute to datacenter networking design and troubleshooting (VLANs, switching, routing fundamentals) where applicable
Build automation and tooling to reduce repetitive manual work across the infrastructure stack
Participate in on-call rotation and incident response, driving root-cause analysis for infrastructure issues
Document architecture, runbooks, and operational procedures
Participate in maintenance activities, platform upgrades, and infrastructure migrations across global data center locations How AI Fits Into This Role We expect you to actively use AI tools to augment (not replace) the skills above.
This looks like:
Scripting & automation: Using AI coding assistants (e.g., Claude Code, Copilot) to draft and refine automation for provisioning, patching, and cluster operations; reviewing and validating output rather than blindly trusting it
Kubernetes operations: Using AI to help generate and audit services and troubleshoot pod/network/storage issues faster by summarizing logs and correlating events across clusters
Incident response: Using AI to accelerate log analysis, correlate symptoms across systems (compute, storage, network) during outages, and draft initial incident timelines/postmortems
Capacity & performance analysis: Using AI-assisted data analysis to identify SAN storage trends, VM resource pressure, or cluster bottlenecks before they become incidents
Documentation:
Using AI to turn tribal knowledge into clear runbooks, architecture diagrams, and onboarding docs, keeping documentation current with less manual overhead
Knowledge gaps: Using AI as a first-pass research tool for unfamiliar storage arrays, network gear, or Kubernetes add-ons, then validating against vendor docs and testing in a safe environment We care about outcomes and judgment, not tool usage for its own sake. You should be comfortable explaining why an AI-suggested command, config, or fix is correct before running it in production. Requirements 5+ years of hands-on experience in Linux systems administration
Strong production experience with Kubernetes (deployment, operations, troubleshooting at scale)
Solid experience with kubevirt-based virtualization
Comfortable using AI tools (coding assistants, LLM-based troubleshooting/research) as part of daily engineering work
Proven experience operating production infrastructure supporting critical internal business services with defined uptime, performance, and recovery objectives. Preferred SAN storage management experience with Pure Storage, Dell, HPE, and/or NetApp platforms
Datacenter networking experience (switching, routing, VLANs)
Experience building internal tooling or automation that incorporates AI/LLM APIs What Success Looks Like Expected to become an active contributor to production infrastructure operations within the first 60-90 days, including participation in incident response, operational reviews, and infrastructure change management
Operate across multiple technology domains.
Incidents are resolved faster because AI-assisted analysis shortens the diagnostic loop
Manual, repetitive operational work steadily decreases as automation (AI-assisted or otherwise) takes it over
The team's institutional knowledge is captured in living documentation rather than a single person's head #LI-RB1 Disclosure Pursuant to Applicable State Equal Pay Transparency Laws - This position has a starting pay range as listed below. Actual salary depends upon many factors, including a candidate's skills, qualifications and experience, location, and salary expectations, and therefore a starting salary at the low end, high end, or even above the stated range may be offered. This position may also be eligible for bonus compensation, commission (if in a sales function), and/or equity grants. Additionally, full-time employees are eligible to participate in our comprehensive benefits framework, including health and wellness benefits, vacation, paid holidays and refresh days, 401(k) retirement plan, life and disability insurance coverages, and other benefits the Company may offer from time to time.
Pay Range :
$101,320.00-$126,650.00 The compensation noted above is based on an annualized hourly rate assuming normal full-time employment. Data Privacy Notice for