Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

STEM Solutions & Consultants LLC

Site Reliability Engineer

Job Description

Job Requirements Tysons, VA Top Secret/SCI Polygraph Unspecified Career Level not specified Salary not specified Join Premium to unlock estimated salaries
Job Description Site Reliability Engineer Description:
We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure. We are proactively building a pipeline of qualified candidates for future opportunities that may arise. If your background aligns with our anticipated needs, a member of our Talent Acquisition team may reach out should a role become available or to proactively screen you for the role.
Responsibilities:
The responsibilities include, but are not limited to:
Monitor and Manage Kubernetes Clusters:
Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on
Kubernetes Kubernetes Management:
Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance
Containerization & Deployment:
Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments
Cloud Infrastructure Management:
Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)
Monitoring & Incident Response:
Set up monitoring solutions, define alerts, an manage the incident response process for any issues related to Jenkins or Kubernetes clusters
Automate Infrastructure Processes:
Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent
Collaborate Across Teams:
Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructure
Security & Compliance:
Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanning
Requirements:
Active
TOP SECRET
clearance or higher is required Bachelor's degree in Computer Science or related field A minimum of two (2) years of experience working with on-premise and off-premise cloud environments Experience with AWS and/or Azure Hands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, K8s, Terraform, Helm, PostgreSQL, or similar technologies Ability to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScript Experience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn) Proactive approach to identifying problems, performance bottlenecks, and areas for improvement Ability to lead and work independently in an Agile/Scrum environment Real passion for developing team-oriented solutions to complex engineering problems Thrive in an autonomous, empowering and exciting environment Great verbal and written communication skills to collaborate multi-functionally and improve scalability Interest in committing to a fun, friendly, expansive, and intellectually stimulating environment
Desired Skills:
Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud Services Experience with deep learning, natural language processing, computer vision, or reinforcement learning Conveys highly technical concepts and information in written form to technical and non-technical audiences The ability to work on multiple concurrent projects is essential. Strong self-motivation and the ability to work with minimal supervision Must be a team-oriented individual, energetic, result & delivery oriented, with a keen interest on quality and the ability to meet deadlines group id: 91130387 Apply now

Benefits

  • Dental Insurance

Career Insights for Site Reliability Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on Virginia data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.

$118,618 / year median in Virginia

Explore Career