Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

TEKsystems c/o Allegis Group

Site Reliability Engineer

Job Description

Job Requirements Naval Base, VA Jb Pearl Harbor Hickam, HI Jacksonville Naval Air Sta, FL San Diego, CA Secret Polygraph not specified Mid Level Career (5+ yrs experience) $85,000 - $117,500 Job Description The SRE will also develop and execute tests focused on system resilience, performance under load, and failure scenarios. They will work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs). Your work will contribute to the development of robust and scalable services that operate reliably in production. Your responsibilities will include maintaining complex computer systems by writing code to automate software releases, monitor systems, and detect and fix problems before users even know there is an issue. You will use these skills to improve site performance and overall reliability. The
SRE-IDAM
role is responsible for supporting, migrating, automation and optimization of software development and deployment process, infrastructure as code, and contribute to the overall maturity of the Site Reliability Engineering program as well as supporting the creation, maintenance, update, modernization, and refresh the capabilities and components of the Navy's Enterprise Network. The
SRE-IDAM
resource provides technical leadership and knowledge of related tasks and coordinates with project managers, customers, stakeholders, and engineers to support ongoing activities as well as new projects to maintain, transform, and modernize the Navy Enterprise Network.
  • Test, maintain (patching, STIGing, and upgrading), troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager (MIM), Active Directory Federation Services (ADFS), DHCP, DNS, WINS, GPOs & PKI.
  • Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform. In addition, as you discover and document system bugs, you have the motivation to go off and fix them yourself.
  • Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines.
  • Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools
  • Identify performance bottlenecks and optimize the performance of cloud infrastructure
  • Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code.
  • Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies.
  • Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
  • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic. Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components. Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals.
  • Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
  • Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
  • Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks.
  • Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management.
  • Must be willing to work nights, weekends, and provide on-call support as needed.
Basic Qualifications
BS + 2-4
years of experience (or MS +

Career Insights for Site Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Virginia data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.

$118,618 / year median in Virginia

Explore Career