Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
AI
Amazon.com, Inc.
Systems Engineer, AWS Incident Response
Job Description
Description As a Systems Engineer on AWS Incident Response, you will be on the front line of AWS incident response. You will lead high-severity calls, triage complex failures across distributed systems, coordinate resolver teams, and drive incidents to mitigation in real time while millions of customers depend on the outcome. Between incidents you will obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice, and you will drive operational improvements that make the incident management ecosystem faster and more accurate. You will make real-time decisions under pressure, deep-diving the largest and most complex technical environment in the world. You will develop expertise across AWS services, networking, and infrastructure, building a breadth of knowledge that few roles offer. Your scope spans all of AWS rather than a single service. You will own operational processes end to end and use data to find the next improvement in how we detect and mitigate faster. You will also have the opportunity to grow your development skills by taking on coding projects that accelerate incident response and reduce toil. This role includes participation in an on-call rotation covering weekdays, weekends, and holidays. On-call shifts fall within your local daytime hours. Key job responsibilities
- Lead high-severity incident response calls end to end: assess impact, coordinate resolvers across AWS service teams, communicate clearly under pressure, manage escalations, and drive the incident to mitigation with documentation throughout.
- Own and run operational health reviews, and build and maintain the dashboards, metrics, and monitoring that surface trends before they become incidents.
- Improve detection accuracy and speed. Identify patterns across events and build proactive mechanisms that prevent recurrence.
- Deep-dive operational data to find systemic issues, measure response effectiveness, and prioritize improvements against what the data shows is degrading.
- Identify gaps in operational processes, documentation, and tooling, and build or improve mechanisms that reduce time-to-detection and time-to-mitigation.
- Apply scripting, automation, and generative AI to accelerate incident response and reduce toil, including where AI can augment human judgment during an incident or surface insight from operational data at scale.
- Work with service teams so that learnings from each incident drive corrective actions to completion, closing the loop between what broke and what gets fixed.
- Mentor peers in your areas of technical and operational strength.
- 3+ years of systems engineering, or 3+ years of technical support experience
- Experience in written and verbal communication skills to communicate with technical and non-technical audiences, including senior leadership
- Experience scripting in one or more language (e.g. Bash, Python, Perl, Ruby), or experience scripting in modern programming languages
- Understanding of operating systems (Linux), networking fundamentals, and distributed systems
- Experience with operational monitoring, alerting, and metrics (CloudWatch, Datadog, Grafana, or equivalent)
- Demonstrated ability to troubleshoot complex technical problems spanning multiple systems or services Preferred Qualifications
- Familiarity with incident management tooling and workflows in a large-scale production environment
- Experience with AWS services and cloud infrastructure
- Experience using generative AI or automation to solve operational problems or accelerate workflows
- Track record of authoring post-incident analyses (post-mortems) and driving corrective actions to completion
- Experience building operational dashboards, runbooks, or automation that improved team efficiency
- Knowledge of infrastructure-as-code and deployment tooling, such as CDK, CloudFormation, Terraform, Ansible, or similar
- Experience coordinating across globally distributed teams and time zones
- Comfort operating with ambiguity and incomplete information, and a self-starting approach to identifying what needs doing Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
- 104,500.00
- 160,000.
Benefits
- Paid Time Off (PTO)
- 401(k) Plans
- Mental Health
- Health and Wellness Programs
Career Insights for Systems Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Washington data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Systems Engineer creates computer and data communication networks for companies and organizations. Plans and designs layout for a network, determines the hardware needed and placement of computers, servers, cables and routers; determines data storage, system capacity, and speed.
$140,932 / year median in Washington
+4% projected growth