We are looking for an experienced AWS Data Engineer to design, build, and optimize scalable batch and real-time data pipelines that process high-volume data, preferably within the Communications domain (e.g., Call Detail Records, network telemetry). The ideal candidate has strong hands-on expertise across Spark, Java, Kafka, and the AWS/Cloudera data lakehouse ecosystem, and is comfortable owning pipelines end-to-end from ingestion and transformation to performance tuning and compliance. Key Responsibilities Design, build, and optimize scalable batch and real-time data pipelines using Apache Spark (Scala/Python), Core Java, and Apache Kafka to process high-volume data. Manage and maintain enterprise data lakehouse architectures spanning AWS and Cloudera environments, leveraging S3, Apache Iceberg, and Hive for scalable storage and querying. Implement Change Data Capture (CDC) and write complex SQL to integrate data from relational databases (Oracle, SQL Server) into dimensional data warehouses. Automate and orchestrate pipeline workflows using Unix shell scripting and Python. Lead and support migration of legacy Hadoop workloads to AWS cloud-native architectures. Monitor, troubleshoot, and tune JVM and Spark job performance to ensure reliability and efficiency at scale. Ensure all data pipelines and storage practices adhere to regulatory compliance and data governance standards. Collaborate with cross-functional teams (data architects, analysts, platform engineers) to translate business requirements into robust data solutions.
Required Skills & Technology:
Strong programming experience in Scala/Python (Spark) and Core Java Hands-on experience with Apache Kafka for streaming data pipelines Working knowledge of AWS services (S3) and Cloudera platform tools (Hive, Iceberg) Strong SQL skills, including CDC concepts, across Oracle and SQL Server Experience with dimensional data warehouse modeling Proficiency in Unix shell scripting and Python for automation Experience migrating on-prem Hadoop workloads to cloud platforms Understanding of JVM internals and Spark performance tuning techniques
Preferred Qualifications:
Prior experience in the Communications/Telecom domain (CDRs, network telemetry data) Familiarity with data governance and regulatory compliance frameworks Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience Minimum years of experience 5-8 years