Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Ztek Consulting

SRE

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
99
out of 100
Average of individual scores

Were these scores useful?

Job Description

First Job Previous Job

1,932 of 10,000

Next Job Last Job Share

More Like This

Salary Not Available SRE

Ztek Consulting

Occupation:

Total other occupations

Location:

Woonsocket, RI

  • 02895
Job Type:

Contract, Full Time (30 Hours or More)

Posted:

09/14/2026

Positions available: 1

Source:

Dice.com

Web Site:

www.dice.com

Job #: jo_2020370614

Job Requirements and Properties

Help for Job Requirements and Properties. Opens a new window. Work Onsite

Full Time Schedule

Full Time Job Type

Contract

Job Description

Help for Partial Job Description. Opens a new window. Skills

  • Site Reliability Engineer
  • Incident Management
  • P2 Incident
  • Kubernetes
  • Python
  • Java
  • GCP
  • P1 Incident
  • Site Reliability Engineer
  • Incident Management
  • P2 Incident
  • Kubernetes
  • Python
  • Java
  • GCP
  • P1 Incident
  • Summary Role
  • SRE
    Remote
  • Woonsocket, RI hybrid Job Description/ Responsibilities
    8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
    Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents structured leadership updates, not just participant involvement
    Experience tuning and validating time-series anomaly detection models in a production observability context this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
    Strong programming proficiency in Python, React, and Java at production quality capable of writing operational tooling that other engineers will rely on
    Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services
    Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
    Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
    Strong cloud platform expertise in Google Cloud Platform (Google Cloud Platform) and Rancher K3s.

Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.

Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal.

Preferred Qualifications

Experience owning Production Readiness Reviews or service launch gates.

Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.

Hands-on chaos or fault injection experience.

TIC (Technical Incident Commander) certification or equivalent structured incident command training

Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact

LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) design or implementation experience

Experience with streaming data platforms: Kafka.

Experience with service mesh and traffic management: Istio, Envoy.

Infrastructure-as-code proficiency at production scale: Terraform or Ansible Employers have access to artificial intelligence language tools ("AI") that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.