Find Jobs
Find Jobs Near You – Available Work in Your Location
DFS Tower Lead
Career Insights for Site Reliability Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Site Reliability Engineer is responsible for designing, implementing, and maintaining highly reliable and scalable software systems and infrastructure. They emphasize automation, code-driven infrastructure, and the use of software tools to manage systems efficiently. Monitors performance within production environments, identifies causes of incidents, and implements preventative measures to ensure software reliability.
$155,100 / year median in California
Job Description
DFS Tower Lead Must Have Technical/Functional Skills AWS, S3, Athena, CloudWatch, Airflow, Databricks, Dremio, Python, SQL, GitHub, metadata, profiling, schema, data quality checks, regression testing, runtime upgrades and RCA fixes Roles & Responsibilities We are seeking an experienced L3 Production Support Engineer to provide advanced support for enterprise data platforms and critical data pipelines. The role requires expertise in AWS, Databricks, Airflow, Dremio, Python, SQL, and data quality frameworks, with a strong focus on incident management, RCA, platform stability, runtime upgrades, and continuous service improvement. Key Responsibilities 1. Provide L3 production support for enterprise data platforms and ensure high availability of business-critical data pipelines. 2. Monitor and support data workloads running on AWS services (S3, Athena, CloudWatch), Databricks, and Dremio environments. 3. Troubleshoot and resolve complex production incidents, ensuring compliance with defined SLAs and minimizing business impact. 4. Perform detailed Root Cause Analysis (RCA) for recurring issues and implement permanent corrective actions. 5. Analyze, debug, and optimize ETL/ELT workflows orchestrated through Apache Airflow. 6. Support and maintain Databricks clusters, jobs, notebooks, workflows, and runtime configurations. 7. Develop and execute SQL and Python scripts for issue investigation, data validation, and operational support activities. 8. Monitor infrastructure and application health using AWS CloudWatch, identifying performance bottlenecks and system anomalies. 9. Manage and support data residing in AWS S3 and queried through Athena, ensuring reliability and performance. 10. Perform data profiling and metadata validation to ensure alignment with business and technical requirements. 11. Implement and monitor data quality checks, reconciliation processes, and validation frameworks to maintain data integrity. 12. Support schema changes, column additions, and data model enhancements while ensuring downstream application compatibility. 13. Execute and validate regression testing during code deployments, platform changes, and infrastructure upgrades. 14. Manage Databricks runtime upgrades, platform patching, and environment maintenance activities with minimal downtime. 15. Utilize GitHub for source code management, version control, release tracking, and deployment support. 16. Work closely with Data Engineering, Architecture, Business, and Infrastructure teams to resolve production issues and improve platform stability. 17. Participate in on-call rotations and provide support for critical incidents during off-hours and planned maintenance windows. 18. Develop operational runbooks, knowledge art icles, and support documentation to improve support efficiency. Salary Range- $110,000-$140,000 a year