Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Insight Global

Site Reliability Engineer

Career Insights for Site Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Pennsylvania data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.

$111,639 / year median in Pennsylvania

Explore Career

Job Description

Insight Global Keyword Location Search view all jobs Site Reliability Engineer Newtown Square, PA Posted 19 days ago Apply Now Job Description SAP NS2 is seeking a Site Reliability Engineer - OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services. This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment. The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments. General Responsibilities Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environment Architect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground up Administer and optimize OpenSearch environments for high availability, resiliency, scalability, security, and performance Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization Analyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layers Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement Maintain the integrity and security of servers, systems, and OpenSearch platform infrastructure Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery Develop and maintain monitoring policies, alerting standards, operational runbooks, and support procedures Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services Ensure proper resource allocation and capacity planning across compute, memory, storage, and network resources Partner with product development and engineering teams to design and enhance service reliability and operational readiness Develop and implement testing strategies and document results for platform changes and operational improvements Support log ingestion, index management, retention policies, lifecycle management, and search performance tuning Work in a diverse environment and cross-train with other global team members Participate in an on-call rotation and support weekend or after-hours operational needs as required We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.

To learn more about how we collect, keep, and process your private information, please review
Insight Global's Workforce Privacy Policy:
https://insightglobal.com/workforce-privacy-policy/. Skills and Requirements 8+ years of hands-on experience in SRE or similar DevOps role Experience deploying, and operating OpenSearch in a Kubernetes-based environment Experience with Elasticsearch (as back up) Expertise with Git Strong background as a Backend SRE / Platform Engineer supporting large-scale distributed systems Ability to build and scale OpenSearch infrastructure from initial deployment through production operations ("0 to 100") Experience with Kubernetes administration, containerized workloads, and cloud-native architectures Experience infrastructure automation, CI/CD, and SRE best practices Apply Now Active Filters Site Reliability Engineer Newtown Square, PA Clear All Powered By Cookie Policy We use cookies to improve your experience on our site. To find out more, read our privacy policy Accept Cookies Decline Cookies