Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
RH
Robert Half
Site Reliability Engineer (SRE)
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Ohio data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$107,188 / year median in Ohio
Job Description
We are looking for an experienced Site Reliability Engineer (SRE) to strengthen observability and operational resilience across a Microsoft Azure environment. This long-term Contract role will work closely with DevOps and engineering teams to establish monitoring standards, expand telemetry coverage, and improve service reliability across cloud-based platforms. The ideal candidate brings deep expertise in Azure operations, modern observability tooling, and production support, with the ability to turn data into actionable insight for faster troubleshooting and stronger system performance.
Responsibilities:
- Create and advance an observability framework for Azure-hosted systems and integrated third-party platforms, ensuring scalable monitoring coverage.
- Develop meaningful dashboards, alerting rules, log analysis views, and distributed tracing to provide actionable insight into application and infrastructure behavior.
- Utilize Azure services such as Azure Monitor, Log Analytics, Application Insights, Managed Prometheus, and Azure Managed Grafana to expand end-to-end visibility.
- Work alongside DevOps and software engineering teams to strengthen platform stability, incident response readiness, and service performance.
- Assess existing monitoring practices to uncover blind spots, reduce unnecessary alert volume, and support quicker root-cause identification.
- Improve insight into the health of applications, infrastructure components, and dependent services across production environments.
- Support reliability-focused engineering efforts by applying SRE principles such as service measurement, alert strategy refinement, and operational readiness improvements.