Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

intone Inc

Site Reliability Engineer

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
99
out of 100
Average of individual scores

Were these scores useful?

Job Description

We are seeking an experienced Site Reliability Engineer to support highly available, cloud-native applications within a Digital Commerce environment. This is a 6+ month contract position based in Eagan, Minnesota. You will focus on monitoring, observability, reliability, and performance of Kubernetes-based applications and services running within Microsoft Azure and Azure Kubernetes Service (AKS), partnering closely with DevOps, Platform Engineering, Production Support, Infrastructure, Network, Security, Architecture, and Development teams. Responsibilities Monitor and improve the reliability, availability, latency, and performance of Kubernetes-based applications and services running on Azure Kubernetes Service (AKS). Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud-based products and services. Design and maintain monitoring and observability frameworks using tools such as Dynatrace, Azure Monitor, and Application Insights. Develop monitoring and alerting that proactively identifies symptoms and performance degradation rather than simply reporting outages. Analyze telemetry, logs, metrics, and monitoring data to identify application and infrastructure bottlenecks. Recognize and address substandard application or infrastructure performance based on established KPIs. Troubleshoot complex issues across distributed, cloud-native systems. Continuously automate and improve capabilities to increase reliability, scalability, performance, and security. Partner with DevOps and Platform Engineering teams to integrate monitoring and observability capabilities into applications, infrastructure, and automated pipelines. Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build and maintain highly available Azure environments. Support incident response and help identify root causes of production reliability and performance issues. Design and configure proactive alerting mechanisms and thresholds to enable rapid identification and resolution of issues. Contribute to the design and implementation of CI/CD pipelines and automated operational processes. Document processes, technical designs, operational procedures, and monitoring standards. Identify cross-team operational risks and drive issues toward resolution through engineering, troubleshooting, and operational improvements. Participate in regulatory and compliance activities as needed. Qualifications Required 5+ years of experience in Software Engineering, Systems Engineering, Operations Engineering, DevOps, SRE, or related technical roles. 2+ years of hands-on Site Reliability Engineering, DevOps, or similar cloud-native engineering experience. Strong experience supporting cloud-native applications hosted within Microsoft Azure. Hands-on experience supporting and monitoring Kubernetes / Azure Kubernetes Service (AKS) environments. Experience monitoring application availability, uptime, latency, infrastructure, and performance across large distributed systems. Strong knowledge of observability and application performance monitoring concepts. Experience troubleshooting complex cloud, application, infrastructure, and system-related issues. Strong debugging and problem-solving skills within distributed environments. Experience designing or implementing CI/CD pipelines. Experience with version control systems such as Git. Working knowledge across systems, networking, security, databases, storage, and cloud infrastructure. Experience collaborating across DevOps, Platform Engineering, Production Support, Infrastructure, Architecture, and Development teams. Strong written and verbal communication skills with the ability to communicate technical monitoring and reliability insights to both technical and non-technical stakeholders. Preferred Deep experience monitoring Kubernetes / AKS environments, containerized applications, services, and workloads. Hands-on experience with Dynatrace. Experience with Azure Monitor and Application Insights. Experience designing and implementing enterprise monitoring and observability frameworks. Experience configuring proactive, symptom-based alerting and thresholds. Experience analyzing telemetry and monitoring

Benefits

  • Dental Insurance