Job Description We are seeking a Network Monitoring Platform Engineer to support and enhance a highly customized network monitoring and service assurance platform responsible for collecting device health and performance data across large-scale environments.
This role will focus on platform enhancements, upgrades, troubleshooting, service migrations, and ongoing operational support. The ideal candidate will have experience supporting monitoring, observability, NOC, or service assurance platforms and possess strong Linux, Python, and automation experience.
Responsibilities-Enhance, tune, and maintain a large-scale network monitoring and service assurance platform-Develop and maintain Python-based services, integrations, and automation workflows-Troubleshoot production issues and provide operational support for platform services-Configure, deploy, upgrade, and maintain platform instances and environments-Build and support custom
REST API
integrations and services-Implement and maintain monitoring and observability solutions using Prometheus and Grafana-Automate deployments and platform operations using Ansible, Bash, Jenkins, and CI/CD pipelines-Perform application and platform upgrades, migrations, and modernization efforts-Troubleshoot and resolve service, performance, and data collection issues-Support Linux-based infrastructure running on Red Hat Enterprise Linux-Work with relational databases to support platform services and integrations-Configure and maintain services responsible for collecting network and device health data-Modify and enhance existing services and applications to support evolving operational standards and requirements Skills and Requirements
- 5+ years supporting network monitoring, observability, service assurance, NOC, or telecommunications platforms-Strong Python development experience (required
- Experience building and consuming REST APIs-Strong Linux administration experience, preferably Red Hat Enterprise Linux-Experience with Prometheus and Grafana or alternative tool-Experience with Ansible, Bash scripting, and CI/CD pipelines using Jenkins (alternative tool ok
- Experience troubleshooting and supporting applications in production environments-Experience deploying, configuring, upgrading, and maintaining platform services-Understanding of networking fundamentals and how device health and telemetry data are collected and monitored-Experience working with relational databases-Experience with Oracle Unified Assurance (Assure1), Netcool, SevOne, ScienceLogic, SolarWinds, LogicMonitor, Splunk, Dynatrace, or similar monitoring platforms-Telecommunications, NOC, or OSS/BSS experience-Java development experience-Experience with fault management, event correlation, and service assurance concepts We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day.
We are an equal employment opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment without regard to race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or the recruiting process, please send a request to HR@insightglobal.com.