A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
Data Center/Cloud and Automation SME Tata Consultancy Services - 3.9 Culver City, CA Job Details $140,000 - $160,000 a year 13 hours ago Qualifications Cloud identity and access management (IAM) Content creation for technical audiences Defect resolution root cause analysis IT user and group management Linux support Load balancers Ansible Procedural guides Infrastructure as Code (IaC) Preventive action implementation IT system monitoring Configuration management Corrective and preventive actions (CAPA) Operating system updates Application deployment Bash System maintenance Reducing cloud infrastructure costs Patch management Cloud service support Continuous improvement Data Security (Data management) Incident management operations support Version control systems Terraform Computer networking Software documentation Scalability System deployment Problem management
Full Job Description Must Have Technical/Functional Skills Linux Administration:
Strong hands on experience with RHEL / Amazon Linux / Ubuntu, including user management, patching, troubleshooting, shell scripting, and performance monitoring.
AWS Cloud:
Practical knowledge of core AWS services such as EC2, VPC, Subnets, Certificate Manager,IAM,S3, Load Balancers, Auto Scaling, and CloudWatch, with security and cost optimization awareness. Infrastructure as Code (Terraform): Experience in writing and managing Terraform modules, variables, and state files to provision and maintain AWS infrastructure. Configuration Management (Ansible): Ability to create and manage Ansible playbooks and roles for OS configuration, automation, and deployment tasks.
Version Control:
Working knowledge of Git for code versioning and collaboration. Understanding of networking fundamentals
Automation & Scripting:
Proficiency in Bash scripting and automation of repetitive operational tasks.
Monitoring & Troubleshooting:
Experience in system and cloud monitoring, incident handling, and root cause analysis.
Security & Compliance:
Understanding of access control, encryption, patch management, and secure configuration practices. Manage and support Linux servers, including installation, patching, monitoring, and troubleshooting. Provision, configure, and maintain AWS infrastructure ensuring availability, security, and scalability. Develop and maintain Infrastructure as Code using Terraform for consistent cloud deployments. Automate system configuration, deployments, and operational tasks using Ansible. Monitor systems and cloud resources; handle incidents, changes, and service requests. Perform root cause analysis and implement preventive actions Enforce security best practices across OS and cloud environments. Collaborate with cross functional teams to support application deployments and platform stability. Maintain documentation, SOPs, and runbooks Manage code/version control through GitHub Continuously improve automation, reliability, and operational efficiency. Roles & Responsibilities Manage and support Linux servers including installation, patching, user management, performance monitoring, and troubleshooting. Design, provision, and maintain AWS infrastructure (EC2, Certificate Manager, Security Groups, VPC, IAM, S3, Load Balancers) ensuring security, availability, and cost efficiency. Implement Infrastructure as Code using Terraform to build, update, and manage AWS resou rces in a consistent and reusable manner. Automate system configuration, deployments, and patching using Ansible playbooks and roles. Code version control through Github. Monitor systems and cloud resources, respond to incidents, and perform root cause analysis. Follow security best practices, access controls, and compliance requirements across OS and cloud platforms. Collaborate with application, network, and security teams to support deployments and changes. Maintain documentation, SOPs, and continuously improve automation and operational efficiency. restructure components.