Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
SI
Slalom, Inc.
Cloud Engineer/Site Reliability Architect
Job Description
Description and Requirements Job Description This is a Kubernetes-heavy platform engineering role supporting large-scale, multi-cloud GPU infrastructure. Success requires deep operational judgment, disciplined debugging, strong automation skills, and the ability to protect platform stability while capacity grows rapidly. What You'll Do
- Operate Kubernetes platforms at significant scale across providers, including Amazon Elastic Kubernetes Service (EKS), CoreWeave Kubernetes Service (CKS), and Google Kubernetes Engine (GKE).
- Own the Kubernetes cluster lifecycle, including node-pool design and management, scheduler troubleshooting, container network interface (CNI) and network-policy troubleshooting, capacity planning, and safe rolling upgrades across large fleets.
- Develop, review, and maintain infrastructure as code with Terraform, along with Python tooling and automation that improve reliability, repeatability, and operational efficiency.
- Define and maintain service-level indicators (SLIs) and service-level objectives (SLOs); build monitoring and alerting that surface meaningful risks before they affect workloads.
- Debug complex distributed-system failures by forming hypotheses, testing them methodically, and separating temporal correlation from causation.
- Provision high-performance computing (HPC) and GPU infrastructure through the Conveyor CI/CD system across AWS, CoreWeave, Google Cloud Platform (GCP), and Oracle Cloud Infrastructure (OCI), with additional providers as the platform expands.
- Coordinate daily with Networking, Storage, Security, and AI/ML platform teams to resolve cross-system dependencies and improve the end-to-end developer and researcher experience.
- Bring forward well-reasoned proposals, solutions, and informed opinions; use available tools and resources to self-unblock and drive issues to resolution.
A proactive working style:
you arrive with proposals, formulate solutions, and communicate informed technical opinions. Clear communication and effective collaboration across engineering teams and technical stakeholders. Working knowledge of AWS or similar Cloud services relevant to compute and data-intensive platforms, including Amazon EC2, Amazon S3, Amazon EFS, and Amazon FSx for Lustre. Experience designing or operating CI/CD pipelines and automated infrastructure provisioning workflows. Knowledge of cloud networking, storage, security, observability, reliability engineering, and platform governance.Nice-to-have:
Experience using AI coding tools responsibly: you remain accountable for the solution, validate generated code, identify edge cases, and reject unnecessary or incorrect output. About Us Slalom is a fiercely human business and technology consulting company that leads with outcomes to bring more value, in all ways, always. From strategy through delivery, our agile teams across 52 offices in 12 countries partner with clients to co-create powerful customer experiences, modern ways of working, and meaningful impact. What sets us apart? We believe work should be challenging and fulfilling, not perfect, but possible. That's why we prioritize purpose, flexibility, connection, and recognition, so our people can thrive and love what they do, most days. Compensation and Benefits Slalom prides itself on helping team members thrive in their work and life. As a result, Slalom is proud to invest in benefits that include meaningful time off and paid holidays, parental leave, 401(k) with a match, a range of choices for highly subsidized health, dental, & vision coverage, adoption and fertility assistance, and short/long-term disability. We also offer yearly $350 reimbursement account for any well-being-related expenses, as well as discounted home, auto, and pet insurance. Slalom is committed to fair and equitable compensation practices. For this role, we are hiring at the following levels and targeted base pay salary ranges: The target base salary range for Senior Consultant in East Bay, Silicon Valley, and San Francisco is $149,000- 185,000. The target base salary range for Senior Consultant in is $149,000
- 185,000. The target base salary range in San Diego, Los Angeles, Orange County, Seattle, Houston, New Jersey, New York City, Westchester, Boston, Washington DC for Senior Consultant is $136,500
- 169,500. The target base salary range for Senior Consultant in all other US based Slalom locations is $125,000
- 155,500.
Benefits
- Paid Time Off (PTO)
- 401(k) Plans
- Health Insurance
- Dental Insurance
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Georgia data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$117,186 / year median in Georgia