Data Developer / Data Engineer Capgemini
- 3.7 Denver, CO Job Details Contract $34.67
- $54.
18 an hour 2 hours ago Benefits Health insurance Dental insurance Vision insurance Retirement plan Qualifications Athena Commercial use (data warehousing systems) Data Integration (Data management) Spark Apache Hive AI platforms (beyond public GPTs) Bash SQL ETL process automation Spark implementation AWS Glue Airflow Version control systems Query execution time improvement Serverless cloud services Linux DevOps automation Query management Cross-functional collaboration GitLab Python Cross-functional communication Hadoop AWS Lambda Full Job Description Denver, CO, United States (On-site) Contract (6 months) Published 6 hours ago linux data modeling SQL hadoop etl AWS lambda query optimization PySpark/Python Pipeline Optimization The Data Developer works within the ETL and Operations team to design, develop, and maintain data pipelines on AWS that drive analytic solutions from diverse, disparate sources (network telemetry, WiFi platform data, billing, and provisioning), routinely processing billions of records per day. The ideal candidate is a skilled data engineer who solves problems in a highly technical, cross-functional environment, with oversight and support from Lead Developers.
Major Duties and Responsibilities:
Coordinate, build, and manage new data ingests from network, WiFi, billing, and provisioning sources. Implement updates, fixes, and optimizations across a large suite of production ETL jobs. Build and tune Spark SQL and PySpark transformations on AWS EMR at billions-of-records-per-day scale. Write and optimize Athena (Trino/Presto) queries for adhoc analysis, validation, and stakeholder support. Contribute to the migration from legacy orchestration (Step Functions) to Airflow / MWAA DAGs and a centralized Python Job Framework. Participate in GitLab merge request reviews and follow CI/CD deployment workflows (staging to production). Monitor production pipelines, respond to data quality alerts and job failures, and partner with analysts and data scientists on aggregation processes.
Required Qualifications:
Strong SQL
- Spark SQL and Athena (Trino/Presto), including comfort moving logic between both dialects. AWS & orchestration
- EMR, S3, Glue Data Catalog, Athena, Lambda, Step Functions, and Airflow / MWAA. Large-scale data
- billions of records per day on partitioned data lakes (Parquet with ZSTD on S3). Python scripting
- PySpark and orchestration / control scripts. Pipelines & optimization
- building and tuning pipelines end to end (partition pruning, join strategy, shuffle and AQE tuning). Version control & shell
- Git / GitLab (branching, MRs, peer review) and Bash for automation on Linux / EMR. Collaboration & learning
- strong communication with analysts, scientists, and architects; ability to pick up new technologies quickly.
Education & experience: Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent experience; 2+ years of data engineering / ETL experience.
Preferred Qualifications:
Apache Iceberg
- table format, MERGE patterns, and Spark-managed DDL (team is actively adopting). Data modeling
- dimensional modeling, slowly changing dimensions, and medallion (bronze/silver/gold) layering. AI-assisted development
- comfort using approved AI coding assistants (Amazon Q, Kiro, GitLab Duo) to accelerate development, review, and troubleshooting; AI / ML pipeline exposure a plus.
Hadoop / Hive
- familiarity with the Hadoop ecosystem and HiveQL for legacy table definitions and SQL-on-Hadoop concepts. The pay range that the employer in good faith reasonably expects to pay for this position is $34.67/hour
- $54.
18/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis. Tundra Technical Solutions is among North America's leading providers of Staffing and Consulting Services. Our success and our clients' success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.
Unincorporated LA County workers:
we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.