Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

The Copley Consulting Group

Sr Site Reliability Engineer

Career Insights for Site Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Illinois data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.

$119,388 / year median in Illinois

Explore Career

Job Description

Financial Software company
Position:
SRE Engineer Location:
Hybrid remote/
Buffalo Grove/Chicago Comp:
Solid Base, bonus, equity
Benefits:
Comprehensive Health, 401K and Equity plan
  • Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
  • Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
  • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
  • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
  • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.
  • Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure—this is a primary engineering responsibility, not an occasional task.
  • Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
  • Author and troubleshoot Azure DevOps pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.
  • Administer and configure Incident.
IO:
alert routing, notification workflows, Slack and OpsGenie integration, and runbook management—operationalizing what exists today and expanding from there.
  • Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on-call rotation design, escalation policies, incident severity classification, and response playbooks.
  • Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
  • Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.
  • Enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance.
  • Collaborate with the Subsystems Platform Team to translate common needs into self-service observability and incident management capabilities.
  • Build lasting team competency through documentation, training materials, and knowledge-sharing sessions that outlast any individual engagement.
WHAT YOU'LL BRING
  • Core SRE Experience

Benefits

  • 401(k) Plans
  • Dental Insurance