Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
SO
System One
AWS Production Support Engineer
Career Insights for Production Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Alabama data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Production Engineer combines knowledge of manufacturing technology and engineering sciences with management theory. Designs the production steps, defines and monitors resources needed, and evaluates efficiency of the overall process.
$86,342 / year median in Alabama
+4% projected growth
Job Description
AWS Production Support Engineer System One - 3.5 Birmingham, AL Job Details Temp-to-hire $75,000 a year 2 hours ago Benefits Health insurance Dental insurance 401(k) Vision insurance Life insurance Qualifications Cloud Logging Responding to infrastructure failures Infrastructure as Code (IaC) Incident report management Triage IT system monitoring Incident management software Public Cloud Incident Escalation Databases Cloud service support Confluence Terraform Splunk Customer support ticket management Resolving technical support tickets Production monitoring Amazon CloudWatch Application Maintenance DevOps automation Cloud monitoring Escalation handling GitLab Production troubleshooting System performance monitoring
Full Job Description Job Title:
AWS Production Support Engineer Location:
Knoxville, TN Type:
Contract To Hire Compensation:
$75,000.00Work Model:
Onsite - onsiteHours:
40.0Security Clearance:
Not specified Overview Leave placeholder text here for recruiter to input Responsibilities Provide monitoring and production support for AWS hosted applications and infrastructure. Monitor dashboards, alerts, application health, infrastructure performance, scheduled workloads, and operational queues. Respond to production incidents and independently initiate technical triage. Review Splunk logs, dashboards, indexes, and basic SPL queries to identify potential issues. Use OpenTelemetry and Amazon CloudWatch logs and metrics during incident investigation. Troubleshoot application outages, system degradation, failed jobs, message processing issues, and infrastructure alerts. Execute documented recovery procedures, including message or transaction replay when required. Manage and update ServiceNow incidents and coordinate with the appropriate support and engineering teams. Follow established runbooks, Confluence documentation, and knowledge articles to resolve known issues. Escalate unresolved or complex issues to the appropriate application, infrastructure, database, security, or engineering team. Execute approved GitLab/Terraform deployment steps and predefined production changes. Support deployments, platform upgrades, disaster recovery, vulnerability remediation, and patching activities. Maintain clear documentation, incident updates, shift handoffs, and knowledge articles. Work independently during assigned night shift coverage while maintaining appropriate communication and escalation. Requirements 5+ years of hands on AWS production support experience supporting cloud hosted applications and infrastructure. Strong incident triage and troubleshooting skills, including responding to alerts, reviewing logs, identifying probable causes, following runbooks, and escalating appropriately. Working knowledge of Splunk, including basic SPL queries, dashboards, indexes, and log investigation. Familiarity with Amazon CloudWatch and OpenTelemetry for reviewing production logs, metrics, and alerts. Experience supporting 24x7 production environments and the ability to work independently during overnight incidents. Experience with ServiceNow or a comparable incident/ticket management platform. Ability to follow documented recovery procedures, including message replay or similar operational recovery actions. Working knowledge of AWS infrastructure, applications, databases, storage, identity, and monitoring concepts. Familiarity with GitLab, Terraform, and CI/CD processes for executing or monitoring approved production changes. Ability to troubleshoot failed jobs, system degradation, service outages, scheduling failures, access issues, and deployment problems. Strong ownership, urgency, communication, and escalation skills. Ability to learn independently from runbooks, Confluence pages, and knowledge articles rather than relying solely on step by step direction. System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan. System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law. #M- #LI- Ref:
#404-IT PittsburghBenefits
- Medical Leave
- 401(k) Plans
- Health Insurance
- Dental Insurance