Find Jobs
Find Jobs Near You – Available Work in Your Location
Senior Automation & Observability Engineer
Choose a Location
This role is available in multiple locations. Pick one to apply.
Job Description
JR014235
At Ensono, our- Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things
- _!
- We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation. We can
- Do Great Things
- because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose: Honesty, Reliability, Curiosity, Collaboration, and Passion.
- About the role and what you'll be doing:
- We are seeking an experienced IoT / Observability Engineer responsible for monitoring, managing, automating, and optimizing enterprise infrastructure, applications, IoT platforms, and enterprise telemetry ecosystems.
- We want all new Associates to succeed in their roles at Ensono. That's why we've outlined the job requirements below. To be considered for this role, it's important that you meet all Required Qualifications. If you do not meet all of the Preferred Qualifications, we still encourage you to apply.
- Key Responsibilities
- Monitoring & Observability
- + Design, implement, and maintain enterprise monitoring and observability solutions.
- Foak & Enterprise Logging/Telemetry
- + Support onboarding, monitoring, and operational management of FOAK (First Office Application Kit) services and enterprise applications.
- Infrastructure & Platform Monitoring
- Monitor and support: + Linux Servers + Windows Servers + VMware Infrastructure + Citrix VDI Platforms + DNS Services + Proxy Services + Middleware Platforms + Integration Services + Enterprise Applications +
IoT Platforms Additional Responsibilities:
+ Investigate performance issues, recurring alerts, and infrastructure anomalies. + Validate monitoring platform health and monitoring coverage. + Monitor capacity, availability, CPU, memory, storage, and service health metrics. + Support platform upgrades, maintenance, and operational readiness reviews.- Database & Data Management
- + Configure and maintain InfluxDB time-series databases. + Manage data retention policies, performance tuning, and capacity planning. + Develop operational dashboards and reports for infrastructure and application performance insights. + Support telemetry data ingestion, storage optimization, and historical trend analysis.
- Event & Incident Management
- + Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms.
- Instana & APM Operations
- + Administer and support IBM Instana monitoring environments.
- Automation & Scripting
- Develop automation solutions using: + Python + PowerShell + Shell Scripting (Bash) +
VBScript Responsibilities:
+ Automate operational tasks, monitoring deployments, and remediation workflows. + Build reusable automation tools to improve operational efficiency. + Integrate monitoring platforms with enterprise automation frameworks. + Support webhook-based automation and event-driven operational workflows.- Configuration Management & Infrastructure Automation
- Implement Infrastructure as Code (IaC) and automation using: + Ansible +
Puppet Responsibilities:
+ Automate server provisioning and configuration management. + Automate monitoring agent deployment and onboarding. + Maintain automation playbooks and deployment pipelines. + Improve operational consistency and reduce manual efforts across environments.- Knowledge Management
- + Maintain SOPs, runbooks, monitoring procedures, and escalation matrices. + Participate in KT sessions, service onboarding, and operational readiness reviews. + Support service transition, migration, and continuous improvement initiatives. + Maintain observability standards and monitoring documentation.
- Required Skills
- Monitoring & Observability
- + Grafana + IBM Instana + SolarWinds + Telegraf + Prometheus + InfluxDB + OpenTelemetry + Grafana Alloy + APM Monitoring + Event Management + Alert Management + Observability Concepts + SLO/SLA Monitoring
- FOAK & Enterprise Telemetry
- + FOAK (First Office Application Kit) Support + Enterprise Logging & Telemetry (ELT) + Log Aggregation & Correlation + Telemetry Data Analysis + Event Correlation + Application Onboarding + Monitoring Standards & Observability Frameworks
- Infrastructure
- + VMware + Linux Administration + Windows Server + Citrix
VDI + DNS
Services + Proxy Services + Middleware Technologies + Infrastructure Performance Monitoring- Scripting & Programming
- + Python + PowerShell + Shell Scripting (Linux) + VBScript
- Automation Tools
- + Ansible + Puppet + Webhooks + Infrastructure as Code (IaC)
- ITSM & Operations
- + ServiceNow + Incident Management + Problem Management + Change Management + ITIL Framework + Major Incident Management
- Nice to Have
- + Docker + Kubernetes + AWS + Microsoft Azure + Google Cloud Platform (GCP) + Jenkins + GitHub Actions + GitLab
CI/CD + REST
APIs + Microservices Monitoring + DevOps & SRE Practices- Preferred Experience
- + 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering.
- Key Competencies
- + Observability & Monitoring + Infrastructure Automation + Enterprise Telemetry Management + FOAK Application Support + Root Cause Analysis + Problem Solving + Performance Optimization + Service Reliability Engineering (SRE) + Cross-Functional Collaboration + Operational Excellence + Continuous Improvement
- Why Ensono?
- Ensono is a place to make better happen - for our clients and for your career.
Some of our benefits include:
+ Unlimited Paid Days Off + Three health plan options + 401k with company match + Eligibility for dental, vision, short and long-term disability, life and AD&D coverage, and flexible spending accounts + Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement + Paid childbearing and paternal leave + Education Reimbursement, Student Loan Assistance or 529 College Funding + Sabbatical leave + Wellness program + Flexible work schedule As of the date of this posting, a good faith estimate of the current pay scale for this role is $113,000 to $147,000 annually based on a full-time schedule. Please note that placement in the range may vary based on numerous factors including but not limited to skills, experience, internal equity, and business needs. In addition to base salary, other compensation programs, depending on eligibility, include- an annual bonus plan based on company and individual performance
- and an equity grant under our Associate Equity Appreciation Program.
Benefits
- Paid Time Off (PTO)
- Sabbatical Leave
- 401(k) Plans
- Health and Wellness Programs
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on national data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$107,921 / year median in the U.S.