Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Robert Half

Site Reliability Engineer (SRE)

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
96
out of 100
Average of individual scores

Were these scores useful?

Job Description

We are looking for an experienced Site Reliability Engineer to strengthen observability and operational excellence across a Microsoft Azure environment.. This Long-term Contract position will work closely with DevOps and engineering teams to create reliable monitoring practices, improve system insight, and support stable production operations. The role is ideal for someone who can translate reliability goals into practical monitoring solutions, actionable alerts, and measurable service performance improvements.
Responsibilities:
  • Create and evolve an observability framework that supports Azure-based platforms as well as integrated third-party services.
  • Develop meaningful dashboards, alerting rules, log analysis workflows, and distributed tracing capabilities using platforms such as Datadog or Dynatrace.
  • Utilize Azure monitoring services, including Azure Monitor, Log Analytics, Application Insights, Managed Prometheus, and Azure Managed Grafana, to expand operational visibility.
  • Evaluate system behavior across applications, infrastructure, and dependent services to improve performance tracking and service health awareness.
  • Collaborate with DevOps and software engineering teams to strengthen reliability practices, accelerate issue detection, and improve response to production incidents.
  • Analyze current monitoring coverage to uncover blind spots, reduce unnecessary alert volume, and support faster root-cause identification.
  • Support infrastructure and application reliability initiatives within cloud-native environments, including Kubernetes-based deployments and Azure-hosted services.
  • Contribute to implementation efforts tied to CI/CD workflows and infrastructure automation to ensure observability is consistently embedded in delivery processes.