Find Jobs
Find Jobs Near You – Available Work in Your Location
DevOps & Site Reliability Engineer (Digital) Deerfield, IL
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Illinois data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$119,388 / year median in Illinois
Job Description
Must Have Technical/Functional Skills Cloud & Platform Engineering (Expert Level) Deep expertise in Microsoft Azure, including: Compute (VMs, App Services, Azure Container Apps) Containers & Orchestration (AKS, Docker) Networking (VNETs, Private Endpoints, Application Gateway, Load Balancers) Storage, Azure Key Vault, Azure Monitor, Log Analytics Proven experience designing enterprise grade, highly available cloud platforms Strong understanding of hybrid and multi cloud architectures (AWS / Google Cloud Platform exposure preferred) DevOps & Engineering Excellence Advanced experience with Azure DevOps and CI/CD pipeline architecture Infrastructure automation using Terraform (modules, state management, governance) Strong scripting skills (PowerShell, Bash) GitOps concepts, branching strategies, release orchestration Site Reliability Engineering (Leadership Level) Ownership of platform reliability, resiliency, and performance Definition and governance of: SLIs, SLOs, SLAs Error budgets and reliability metrics Advanced observability strategy: Metrics, logs, traces, alerts, dashboards using Dynatrace Incident response leadership, RCA facilitation, and long term remediation planning Experience operating 99.9%-99.99% availability systems Containers, APIs & Integration Leadership-level experience with AKS-based platforms, ingress, and scaling strategies Understanding of microservices, API-led and event-driven architectures Familiarity with Azure Integration Services (Service Bus, Event Hub, API Management) Security, Compliance & Cost Secure cloud design using Key Vault, managed identities, RBAC Cost optimization (FinOps mindset) across cloud infrastructure Roles & Responsibilities Act as Lead SRE for client''s Digital platforms, owning reliability and stability outcomes Define and enforce SRE standards, best practices, and operating models Architect and govern highly available, scalable cloud platforms Lead the design and implementation of CI/CD and IaC strategies Establish proactive monitoring, alerting, and incident prevention mechanisms Own major incident leadership, RCA execution, and corrective action tracking Partner with application, security, and architecture teams to build reliability by design Drive automation to reduce toil and improve operational efficiency Mentor and coach SRE and DevOps engineers across teams Influence roadmap decisions with a reliability, scalability, and cost lens