Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

E-Solutions Inc.

SRE Azure Cloud

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
100
out of 100
Average of individual scores

Were these scores useful?

Job Description

SRE Azure Cloud (pleasanton, CA) | 09/29/26 Job Description SRE Azure Cloud ||

Pleasanton CA Ideal Candidate:

A hands-on Senior SRE with expertise in Stage/Production deployments, Azure-based microservices, observability (Grafana, Prometheus, Graylog, Azure Monitor), troubleshooting using Application Insights, database operations, and customer escalation management, focused on maintaining highly reliable and available production services. Position Summary

We are seeking a highly skilled Senior Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational stability of customer-facing applications and services. This role is focused on Stage/Production deployments, Monitoring & Observability, Troubleshooting, Incident Management, and Customer Escalation support across Azure-based microservices environments.

The ideal candidate will possess strong experience in cloud operations, microservices, production support, deployment automation, observability platforms, and database troubleshooting, with a proven ability to rapidly diagnose and resolve complex production issues.

Key Responsibilities

Production & Stage Operations

Manage and support Stage and Production environments.

Execute application, infrastructure, configuration, and database deployments.

Validate releases, perform health checks, and coordinate rollback activities.

Support change management and production readiness reviews.

Monitoring & Observability

Build and maintain dashboards, alerts, and monitoring solutions.

Monitor application, infrastructure, and database health using logs, metrics, traces, and telemetry.

Improve observability coverage and reduce alert noise.

Proactively identify reliability and performance issues before customer impact.

Troubleshooting & Incident Response

Troubleshoot software, infrastructure, configuration, deployment, and database-related issues.

Lead incident response activities and production recovery efforts.

Perform root cause analysis (RCA) and implement preventive actions.

Develop operational runbooks and troubleshooting documentation.

Customer Escalation Management

Investigate and resolve customer-reported production issues.

Act as a technical lead during high-priority incidents.

Partner with Engineering, Product, and Customer Support teams to drive issue resolution.

Provide timely communication and status updates during major incidents.

Required Technical Skills

Cloud & Infrastructure

Microsoft Azure

Azure Kubernetes Service (AKS)

Azure Virtual Machines

App Services

Azure Storage

Azure Networking

Application Gateway

Azure Key Vault

Microservices & Containerization

Kubernetes

Docker

Helm Charts

Microservices Architecture

REST APIs

Event-Driven Architecture

Distributed Systems Troubleshooting

CI/CD & DevOps

Azure DevOps Pipelines

Bitbucket

Git

Helm-based Deployments

CI/CD Release Management

Deployment Automation

Monitoring & Observability

Grafana

Prometheus

Graylog

Azure Monitor

Application Insights

Log Analytics

Alerting & Dashboard Management

Distributed Tracing

SLI/SLO Monitoring

Troubleshooting Expertise

Application Performance Issues

Production Incident Management

Configuration & Environment Issues

Deployment Failures & Rollbacks

Kubernetes & Container Troubleshooting

Network & Connectivity Issues

Root Cause Analysis (RCA)

Databases

Azure

SQL / SQL

Server

PostgreSQL / MySQL

Cosmos DB

Redis

Query Performance Tuning

Database Monitoring

Backup & Recovery

Automation & Scripting

PowerShell

Python

Bash

Preferred Experience

Supporting enterprise SaaS applications in Production environments.

Azure-based microservices platforms running on AKS.

Customer-facing production support and escalation management.

24x7 on-call and incident response environments.

Site Reliability Engineering (SRE) best practices including SLIs, SLOs, MTTR, and service availability management.

Key Competencies

Strong troubleshooting and analytical skills.

Production support and incident management expertise.

Customer-first mindset.

Excellent communication and stakeholder management.

Ability to perform effectively during critical outages and high-severity incidents.

Continuous improvement and automation mindset.

SRE Azure Cloud1SRE Azure Cloud ContractUnited States