A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
4+ Months (Contract to hire) Overview A successful Site Reliability Engineer combines software engineering expertise with operational excellence to build and maintain highly reliable, scalable, and performant systems. This role focuses on improving service availability through automation, monitoring, and continuous optimization of production environments. Responsibilities Design, develop, and implement automation solutions to streamline operational tasks and system health checks. Monitor and maintain production systems using observability platforms such as Dynatrace, Splunk, or similar tools. Analyze system performance and reliability metrics to identify gaps and drive process improvements. Partner with engineering and operations teams to provide system design guidance, platform support, and capacity planning. Develop and maintain comprehensive documentation, including standard operating procedures (SOPs), configuration details, and infrastructure diagrams. Minimum Qualifications Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. 5+ years of experience in Site Reliability Engineering, preferably within a product-based or financial technology environment. 4+ years of experience in automation and scripting using languages such as Python, Java, Ansible, or PowerShell. 4+ years of experience with monitoring and observability tools such as Dynatrace, Splunk, Grafana, or similar platforms. Preferred Qualifications Experience with CI/CD pipelines and tools such as GitLab, Harness, Terraform, Nexus, or SonarQube. Strong analytical and problem-solving skills, with the ability to perform root cause analysis and implement proactive solutions. Excellent communication and collaboration skills, with the ability to work effectively across cross-functional teams