Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

ABS Group

Sr. Manager, SRE, Operations & Product Support

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
100
out of 100
Average of individual scores

Were these scores useful?

Job Description

ABS Group Digital Solutions is seeking a Lead, Site Reliability Engineering, Operations & Product Support to establish the production reliability and support operating model for its next-generation Fleet Management System (FMS). The role will help take FMS from development through beta, customer migration, and scaled SaaS operations. It combines hands-on reliability engineering with leadership of incident response, operational readiness, technical support escalation, and post-launch improvement. Working with the Platform Engineering & Cloud Architecture lead, application engineering, QA, security, and customer support teams, this person will ensure the product is observable, recoverable, and supportable. The lead will use conventional and AI-assisted automation to streamline operations while establishing clear boundaries between technical SRE ownership and customer-facing support ownership.
What You Will Do:
Define service health measures and reliability objectives for FMS, including availability, latency, error rates, recovery expectations, and appropriate service-level indicators and objectives. Establish end-to-end observability across applications, infrastructure, integrations, and AI-enabled services through actionable logs, metrics, traces, dashboards, health checks, and alerts; assess where AI-assisted anomaly detection can improve signal quality. Lead the technical incident-response model, including severity definitions, on-call and escalation practices, incident coordination, recovery procedures, and post-incident reviews; use AI-assisted summarization and evidence gathering where it improves response without replacing human judgment. Work with engineering and platform teams to design for resilience, performance, capacity, backup and recovery, and safe operation under expected customer and data growth. Define and implement operational-readiness criteria for beta and production releases, including monitoring, runbooks, rollback plans, ownership, support handoffs, and post-release validation. Establish the technical support escalation model and partner with customer support, product, and engineering to resolve issues and turn recurring incidents and tickets into permanent fixes; evaluate AI-assisted ticket categorization and knowledge retrieval to speed technical triage. Support customer migrations, go-lives, and post-launch stabilization by preparing technical monitoring and response plans, triaging production issues, and incorporating lessons into repeatable procedures. Automate routine operational tasks, health checks, deployment verification, incident triage, and recovery; use AI where it demonstrates improved speed or accuracy, with access controls, auditability, and human approval for production-impacting actions. Track reliability, incident, supportability, and operational-efficiency trends; communicate risks, corrective actions, and progress to engineering and program leadership, including evidence of whether AI-assisted workflows reduce toil or improve outcomes. Help build and mentor an SRE/production-operations capability as FMS moves from initial releases to scaled customer use.
What You Will Need:
Education and Experience Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent relevant experience. 10+ years of relevant experience in site reliability engineering, production engineering, cloud operations, or software operations, including experience leading incident response or operational improvement across teams. Demonstrated experience operating production web or SaaS services and improving their reliability through software engineering and automation. Experience establishing observability, on-call practices, runbooks, and service-health or reliability measures for production systems. Experience partnering with software engineering, platform engineering, and customer-facing support teams during releases, incidents, and customer go-lives. Experience applying AI-assisted tools or workflows to technical operations, incident triage, monitoring analysis, support knowledge retrieval, or operational automation, with an understanding of how to validate results before use in production. Experience with Azure cloud preferred. Knowledge, Skills, and Abilities Strong command of SRE practices, including SLIs/SLOs, incident response, root-cause analysis, performance, capacity, resilience, and disaster recovery. Experience operating production SaaS applications on Azure, including compute, networking, identity, storage, containers, databases, integrations, and security. Ability to build effective observability and on-call practices using telemetry, logs, metrics, traces, and tools such as Azure Monitor, Application Insights, and Log Analytics-without creating unnecessary alert noise. Ability to automate secure deployments and operational workflows using scripting, APIs, CI/CD, infrastructure as code, managed identities, and Key Vault. Sound judgment on release risk, rollback, customer impact, and the responsible use of AI-assisted operations, including data protection and human oversight of production-impacting actions. Clear communication and collaborative leadership across engineering, support, and business stakeholders, including the ability to drive improvements without direct ownership of every team or system.
Reporting Relationships:
Reports to the Sr Director, Platform Engineering & Cloud Architecture. The role will initially lead cross-functional operational practices; direct-report scope will be determined as the SRE and production-operations capability scales. Customer-facing support teams retain ownership of routine customer communications and frontline support.