Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

JP Morgan Chase Company

Technology Support Lead - Major Incident Management

Career Insights for Site Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Ohio data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.

$107,188 / year median in Ohio

Explore Career

Job Description

Join our dynamic team to innovate and refine technology operations, impacting the core of our business services. As a Technology Support Lead in Employee Platforms Production Management, you will play a leadership role in ensuring the operational stability, availability, and performance of our production services. Critical thinking while overseeing day-to-day maintenance of the firm's systems will be key and set you up for success as you navigate tasks related to identifying, troubleshooting, and resolving issues to ensure a seamless user experience. Job responsibilities Lead teams of technologists that provide end-to-end application or infrastructure service delivery for the successful business operations of the firm Serve as an senior incident manager for major incidents, driving triage, impact assessment, escalation, stakeholder communications, decision making, and service restoration through closure. Drive post-incident review, root cause analysis, recurrence prevention, and remediation tracking, ensuring actions are measurable and aligned to resiliency, audit, and control expectations. Lead with clarity, integrity, and urgency during time-sensitive situations, balancing technical judgment, risk management, customer impact, and transparent executive-level communication. Leads team adoption of enterprise-authorized AI capabilities within the work environment to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data. Execute policies and procedures that ensure operational stability and availability and lead incident, problem, and change management in support of full stack technology systems, applications, or infrastructure Monitor production environments for anomalies, address issues, drive evolution of utilization of standard observability tools, and escalate and communicate issues and solutions to business and technology stakeholders through incident resolution and service restoration Applies reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations. Required qualifications, capabilities, and skills 5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits and awareness of data sensitivity. Ability to review and validate AI-assisted incident recommendations before action, escalating when uncertain and ensuring outcomes align to operational, security, and auditability expectations. Experience managing applications or infrastructure in a large-scale technology environment both on premises and public cloud Proficient in observability and monitoring tools and techniques Experience executing on processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skills Demonstrated experience leading incident management in a large-scale technology environment, including major incident command, escalation management, restoration coordination, and executive stakeholder communication. Strong working knowledge of ITIL-aligned incident, problem, and change management practices, including post-incident review, root cause analysis, corrective actions, and recurrence prevention. Proven leadership skills, including the ability to coach technologists, align teams around priorities, make timely decisions under pressure, and communicate clearly with senior business and technology stakeholders. Ability to manage operational risk, control requirements, resiliency expectations, and auditability needs while maintaining focus on user impact and service restoration.