Find Jobs Near You – Available Work in Your Location
Technology
Computer Systems Engineer / Architect
Lutherville-Timonium, MD
Find & Apply For Computer Systems Engineer / Architect Jobs in Lutherville-Timonium, Maryland
Browse jobs from a variety of sources below, sorted with the most recently published, nearest to the top. Click the title to view more information and apply online.
4+ years experience SRE
2+ years experience with monitoring and observability platforms (NewRelic, Prometheus, Grafana, etc.)
2+ years Cloud Experience, Chaos Engineering experience preferred
Job Description:
The InfoSec SRE is a pivotal role focused on engineering and advancing the maturity of the organization's site reliability framework development, adoption, and integration. This position establishes the foundational framework for Site Reliability Engineering (SRE) design fabrics within InfoSec. Additionally, the role supports management of runbook repositories, observability, and related technologies, acting as a trusted advisor to peers and stakeholders across the organization.
Partner with InfoSec and engineering teams to define reliability standards and operating models
Establish and drive adoption of SRE principles across teams and provide guidance to onboarding organizations to the framework.
Design and implement incident response processes and escalation models
Integrate and optimize alerting and on-call workflows using PagerDuty
Develop and maintain operational runbooks and playbooks
Lead or support incident reviews and postmortems with a focus on continuous improvement
Integrate SRE workflows with enterprise platforms such as ServiceNow
Automate operational tasks, incident workflows, and reporting
Improve system resilience through automation and self-healing mechanisms
Design and execute chaos engineering experiments to validate system resilience
Identify failure modes and proactively address system weaknesses
Collaborate with engineering teams to improve fault tolerance and recovery strategies
Drive adoption of reliability best practices across the organization
Provide guidance and mentorship on SRE principles Required Skills/Knowledge
4-5 years of experience with SRE, DevOps, or Infrastructure engineering
Skills with multi-cloud environments (AWS, Azure) as well as on-prem integrations
Strong experience with monitoring and observability platforms (NewRelic, Prometheus, Grafana, etc.)
Experience with incident management and on-call systems (e.g. PagerDuty)
Familiarity with ITSM platforms, specifically ServiceNow and integration via API
Solid understanding of system reliability, performance, and scalability
Automation scripting
Runbook generation
Desired Skills/Knowledge:
Experience working within or alongside InfoSec teams
Proficiency with scripting languages such as Python, Bash, Ruby etc.
Knowledge of security or security-adjacent tools, controls, and compliance frameworks
Experience implementing SLO's, SLAs, and error budgeting
Exposure to chaos engineering practices (specifically, SteadyBit as a tool)
Strong communication and collaboration skills, specifically documenting
Systems thinking and problem-solving mindset
A strong focus on automation and continuous improvement methodologies
Data-driven decision making
Ability to operate in ambiguous, greenfield environments
Understanding of Infrastructure as Code
Passion for the work and responsibility to the consumer
Understanding of public cloud platforms and services
Positive attitude with a strong desire to continuously learn and adapt
Talteam Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.