Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
TC
Tata Consultancy Services Limited
SRE Engineer
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on New Jersey data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$125,134 / year median in New Jersey
Job Description
Must Have Technical/Functional Skills
- 6-7 years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations.
- Strong understanding of Linux administration and troubleshooting.
- Hands-on experience with AWS cloud services (EC2, RDS, IAM, VPC, CloudWatch, S3).
- Experience with monitoring, alerting, and observability tools.
- Knowledge of incident management, problem management, and RCA processes.
- Experience with automation and scripting using Shell and/or Python.
- Working knowledge of PostgreSQL and MySQL databases.
- Experience with Git version control.
- Understanding of CI/CD concepts and tools such as Jenkins. Roles & Responsibilities
- Design, implement, and maintain highly available and reliable production systems.
- Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform.
- Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services.
- Perform Linux server administration, troubleshooting, patching, and performance tuning.
- Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk.
- Participate in incident management, root cause analysis (RCA), and problem management activities.
- Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
- Support PostgreSQL and MySQL databases for operational and basic administration tasks.
- Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes.
- Drive automation, observability, capacity planning, security, and operational best practices.
- Participate in on-call support and production issue resolution.