Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

ECS Federal, LLC

Senior Site Reliability Engineer

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
99
out of 100
Average of individual scores

Were these scores useful?

Job Description

Everforth ECS is seeking aSeniorSite Reliability Engineerto work remotely.

Everforth ECS is seeking talented professionals to join our successful and growing team in building the next-generation Continuous Diagnostics and Mitigation (CDM) Cyber data solution. The CDM Program is the Cybersecurity and Infrastructure Security Agency's (CISA) dynamic approach to strengthening the cybersecurity of Federal networks and systems through better awareness and visibility into their security posture and cyber threats. ECS is responsible for designing, building, deploying, operating, and maintaining a complete 'Data Services' solution which includes the collection, normalization, visualization, and sharing of cyber data from more than 100 Federal agencies. The CDM Data Services product is an integrated suite of multiple Commercial Off the Shelf (COTS) products, software configuration packages, and custom code which work together to operate as an integrated solution tailored to meet Department of Homeland Security (DHS) requirements.

We are seeking professionals who thrive in a dynamic, fast-paced, and highly collaborative environment where problem-solving, critical thinking, and a holistic approach to serving the mission are key. Our program operates within the Scaled Agile Framework (SAFe). An aptitude and enthusiasm for continuous learning, improvement, and cyber security is a must!
Role & Responsibilities:
ECS is seeking a talentedSeniorSite Reliability Engineer (SRE)to play a key role indefining, implementing, and growing our SRE practiceto ensure the reliability, availability, and performance of our critical production environments.

The SeniorSREwill contribute toa culture ofcontinuous improvement,identifyingareas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency.

Thesuccessful candidate willhavedemonstratedhands-on experiencedesigning, implementing, andmaintainingsolutions to ensure thatsystems, includinginfrastructure and applications,are resilient,highly available, and performant.

TheSenior SREwillalsoplay a critical role in definingand measuringthe Service Level Objectives (SLOs) and Service Level Indicators (SLIs)for our solution.

TheSeniorSREwillbe responsible forsettingup comprehensivelogging,monitoring,andalertingsolutionsusing the Elasticstackand other toolsas necessaryto ensure the continuous performance of services.

Additionally, they will respond to incidents, perform root cause analyses, and implement solutions to preventreoccurrences.

TheSeniorSRE will work in close collaborationwith other SRE team members,developers,testers,infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.
Salary Range:
$118,000 - $177,000General Description of BenefitsMust be a US citizen withthe abilityto obtain Public Trust Suitability.6+ years of experience as aSite ReliabilityEngineer(SRE)or equivalent6+ years ofdemonstratedexperience designing,implementing, andmaintainingobservabilitysolutionsto include logging, monitoring, and alerting6+ years of hands-on experience withSREtools (e.g.,Elastic, Prometheus, Grafana, Splunk, etc.)3+ yearsdefiningand measuring SLOs and SLIs3+ years of relevant experience using cloud platforms (AWSGovCloud preferred)3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.)Strong knowledge of microservices, containerization, and orchestration tools (Docker,Kubernetes)Proven ability to collaborate with cross-functional teams (development, testing,andproduct)to integrate reliability and observability into the software development lifecycleStrong problem-solving and analytical skillsProactive, detail-oriented approach toidentifyinginefficiencies and implementingimprovements.

Proficient in developing Synthetic monitoring scripts using typescript.

Benefits

  • Dental Insurance