Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
AE
American Express
Site Reliability Engineer I
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Florida data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$107,921 / year median in Florida
Job Description
View More Jobs Site Reliability Engineer I Sunrise, FL, United States (Hybrid) Apply Now Job Description Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability. Responsibilities Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA) Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions Participate and lead the Development change review and change validation processes Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers