Find Jobs
Find Jobs Near You – Available Work in Your Location
Sr. Splunk Observability & APM Integration Engineer/Architect
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on New Jersey data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$125,134 / year median in New Jersey
Job Description
We are seeking a Senior Splunk Observability & APM Integration Engineer/Architect for a 6-month contract position based in Warren, New Jersey. This is a hands-on role focused on building an enterprise monitoring and observability framework from the ground up, establishing Splunk as the central analytics and dashboarding platform while integrating existing monitoring tools, application telemetry, infrastructure data, batch pipelines, and messaging platforms. Responsibilities Build Splunk dashboards, correlation searches, and SPL-based operational views that support a single-pane monitoring model. Integrate LogicMonitor, AutoSys, OpenTelemetry, and custom data collectors into Splunk. Establish practical APM capabilities, including distributed tracing using an OpenTelemetry-first approach. Assess and advise on buy-versus-build decisions across Splunk APM/SignalFx and Dynatrace. Design synthetic monitoring for critical user journeys, including SLOs, alerting, escalation workflows, and PagerDuty routing. Develop automation using shell scripting and PowerShell across Windows and Linux environments. Assess applications to identify meaningful logging, metric, and tracing opportunities. Build tactical dashboards for immediate operational visibility and strategic dashboards for long-term support and incident response needs. Create architecture documentation, implementation patterns, operational runbooks, and knowledge-transfer materials. Develop telemetry solutions for legacy, custom, unconventional, or unsupported platforms. Partner with technical and operational stakeholders to deliver measurable monitoring outcomes. Qualifications Required Advanced Splunk administration experience with expert-level SPL skills. Strong background in log observability, operational monitoring design, and enterprise dashboard development. Hands-on APM, application instrumentation, and distributed tracing experience. Strong knowledge of OpenTelemetry, including Collector, OTLP, telemetry pipelines, and instrumentation patterns. Experience with Splunk APM/SignalFx and Dynatrace, including synthetic monitoring capabilities. Strong shell scripting and PowerShell automation skills. Experience with LogicMonitor and AutoSys, or comparable monitoring and workload automation platforms. Solid understanding of logs, metrics, traces, SLOs, alerting, incident response, and modern observability practices. Ability to instrument difficult-to-monitor platforms and normalize telemetry into a unified observability framework. Strong application assessment skills, including identifying logging opportunities and translating findings into dashboards and monitoring strategies. Clear verbal and written communication skills with a collaborative approach to documentation and knowledge transfer. Ability to work effectively with engineering, infrastructure, operations, and support stakeholders. Preferred Experience with PagerDuty. Experience with fintech platforms. Experience in regulated environments. Experience with enterprise incident-response dashboards.