Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
C
CACI
Network Management Systems (NMS) Operations Tier 3
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on New Mexico data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$113,710 / year median in New Mexico
Job Description
Job Title:
Network Management Systems (NMS) Operations Tier 3Job Category:
Information Technology Time Type:
Full time Minimum Clearance Required toStart:
TS/SCI withPolygraph Employee Type:
Regular Percentage of Travel Required:
Up to 10%Type of Travel:
Continental US •The Opportunity:
We are seeking an experienced Network Management Systems (NMS) Engineer to oversee the operations of the production NMS tools suite. The ideal candidate will have advanced knowledge of network management systems and be responsible for monitoring, maintaining, and optimizing our production Linux based NMS infrastructure.Responsibilities:
- Administer, configure, and troubleshoot Linux-based systems (e.g., CentOS, Ubuntu, RHEL) in an air gapped environment.
- Monitor, configure, and optimize Linux servers for NMS applications (e.g., Riverbed, SolarWinds, Network Node Manager)
- Monitor system performance, identify bottlenecks, and implement improvements (e.g. Prometheus, collectd, Grafana, InfluxDB).
- Troubleshoot and resolve system issues, including system failures, performance problems, and network-related issues.
- Develop and implement automation scripts to improve system management efficiency
- Work closely with DevOps and engineering teams to identify areas for process improvement and automation.
- Analyze system performance data and provide recommendations for optimization
- Execute projects related to NMS upgrades, migrations, and integrations
- Manage system updates, patches, and security configurations to ensure systems are up-to-date and secure.
- Provide support for automation-related incidents and work on optimizing system health and uptime.
- Mentor junior team members and provide technical guidance
- Collaborate with cross-functional teams to ensure system reliability and security
- Ensure high availability, reliability, and scalability of Linux environments to support the NMS.
- Participate in on-call rotations for critical incident response
Qualifications:
Required:
- Bachelor's degree in Technical field or equivalent work experience
- 10+ years of related work experience
- TS/SCI with Poly required
- Strong knowledge of Linux operating systems (e.g., Red Hat, CentOS, Ubuntu)
- Experience with cloud platforms (AWS, Azure, GCP) and on premise virtualization platforms (VMware, libvirt, KVM) and their monitoring tools
- Proficiency in shell scripting and at least one programming language (e.g., Python, Bash)
- Experience with configuration management tools (e.g., Ansible, Puppet, Chef)
- Expertise in network management tools and platforms (e.g., Riverbed, SolarWinds, Network Node Manager)
- Familiarity with ITIL processes and best practices
- Excellent troubleshooting, problem-solving and analytical skills
- Strong communication and teamwork abilities
Desired:
- Relevant certifications (e.g., RHCE, ITIL)
- Hands-on experience with CI/CD tools like Jenkins, GitLab CI, GitHub Actions, or similar.
- Experience with monitoring tools such as Prometheus, collectd, Grafana, InfluxDB
- Knowledge of log management and analysis tools (e.g., Elastic)
- Understanding of DevOps practices and CI/CD pipelines -