Long Term Contract We are seeking an experienced Data Engineer (70% engineering, 30% analytical) with strong proficiency in modern cloud-native data platforms, large-scale data processing, and the ability to contribute meaningfully to analytical and data science workflows. The ideal candidate combines deep Databricks and PySpark expertise with finance or payroll domain experience and the analytical fluency to partner closely with data scientists on macroeconomic and financial data research. No two days are the same. The Data Engineer will work across production pipeline development, platform optimization, exploratory data analysis, and close collaboration with data scientists, economists, and business stakeholders to build and maintain scalable data platforms for financial analytics and research. To thrive in this role, you'll need to be a dynamic and passionate individual who can move fluidly between engineering and analytical work — building production ETL in the morning and profiling a new dataset in the afternoon. You should have a solid track record in data platform development with finance or payroll domain experience.
WHAT YOU'LL DO
Here's what you can expect on a typical day in the life of a Data Engineer. Our Data Engineers do more than engineering — they combine platform development with analytical exploration to build scalable solutions that power financial and economic research. Experience translating business requirements into scalable data solutions. Data Engineering & Platform Development Design, develop, and maintain scalable data pipelines for ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets. Build and support Databricks-based platforms that enable financial research and analytical workloads. Implement ETL/ELT frameworks using PySpark and Delta Lake for structured and unstructured data from internal and external sources. Develop data models and data marts optimized for analytical, reporting, and machine learning use cases. Ensure data quality, consistency, lineage, governance, and observability across all data assets. Optimize performance for large-scale datasets (billions of records, multi-TB, time-series payroll data). Analytical & Data Science Support Perform exploratory data analysis — profile datasets, identify distributions, outliers, missing patterns, and data drift. Translate data scientist logic into efficient, scalable PySpark implementations (e.g., cross-sectional metrics, time-windowed aggregations). Build validation dashboards and exploratory notebooks to verify data pipeline outputs and data quality. Support feature engineering — implementing complex aggregation and transformation logic at scale. Independently verify and sanity-check analytical outputs — flag when numbers don't make sense. Conduct ad-hoc analytical work using pandas/numpy alongside PySpark for research support. Modern Data Architecture Work within established lakehouse architecture using Databricks and Delta Lake. Contribute to architecture design discussions — understand tradeoffs in catalog design, medallion patterns, and data mesh concepts. Implement CI/CD pipelines using Databricks Asset Bundles (DAB), Bitbucket Pipelines, and Jenkins. Manage Unity Catalog governance, access patterns, and schema design. Ensure security, compliance, and data governance standards are met. AI-Assisted Development Leverage AI coding tools (GitHub Copilot, Amazon Q, Kiro, or equivalent) to accelerate development. Critically review AI-generated code for correctness, performance, and maintainability. Integrate AI-assisted workflows into daily engineering and analytical tasks.
TO SUCCEED IN THIS ROLE, YOU ARE SOMEONE WITH
Required Qualifications:
Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related field. Experience in Data Engineering or Data Platform development. Finance or payroll domain experience — understands payroll data structures, pay period logic, compensation/deduction relationships, or has worked in financial services data environments. Experience handling large-scale datasets (billions of records, multi-TB, time-series data). Strong Databricks proficiency with deep understanding of: Unity Catalog (governance, access patterns, catalog/schema design) Delta Lake internals (optimization, clustering, change data feed, versioning) Databricks Workflows (orchestration, dependencies, monitoring) Databricks Asset Bundles or equivalent deployment frameworks Exposure to architecture-level design decisions (can participate in architecture discussions, understands tradeoffs).
Strong proficiency in:
Python SQL PySpark Data Modeling ETL/ELT Development Analytical fluency: exploratory data analysis, basic statistical concepts (correlation, distributions, time-series patterns), feature engineering support. Proficiency with pandas/numpy for ad-hoc analytical work alongside production PySpark. CI/CD implementation experience (Bitbucket Pipelines, Jenkins, automated deployment). Proficiency with AI-assisted development tools (GitHub Copilot, Amazon Q, Kiro, or equivalent). Data quality and validation frameworks. Preferred Qualifications Experience with macroeconomic, capital markets, or financial services data. Exposure to data architecture design (lakehouse patterns, data mesh concepts, medallion architecture). Experience with streaming/event-driven pipelines (Kafka, Structured Streaming). Census/geographic data processing experience (TIGER, FIPS codes). Infrastructure-as-code familiarity (Terraform, CDK). Databricks certifications (Associate or Professional level). Experience migrating legacy platforms to modern stack (Glue/EMR/HDInsight → Databricks). Experience supporting machine learning and AI-driven analytics solutions. Data visualization experience (Power BI, Tableau, Databricks dashboards). Programming Python, SQL, PySpark, Scala (Preferred) Data Engineering & Analytics Platforms Databricks, Apache Spark, Delta Lake, Unity Catalog, Databricks Asset Bundles, Kafka Databases SQL Server, PostgreSQL, Delta Tables, NoSQL Databases Data Visualization & Reporting Power BI, Databricks Dashboards, Python visualization libraries (matplotlib, plotly) DevOps & Automation Git, Bitbucket Pipelines, Jenkins, Terraform, CI/CD Pipelines, DAB Key Competencies Can move between engineering and analytical mode — builds production ETL and profiles datasets with equal comfort. Strong analytical and problem-solving skills with statistical literacy. Builds reliable, maintainable pipelines — thinks about the next person who will maintain the code. Uses AI tools effectively to accelerate both engineering and analytical work. Enough domain context to question data anomalies and propose data-driven improvements. Can own a notebook from "explore this data" through to "here's what I found." Excellent communication and documentation skills. Strong understanding of data governance, data quality, and metadata management. Ability to work in a fast-paced, data-driven environment.
Nice-to-Have Domain Experience Payroll Data Structures and Processing Macroeconomic Analysis and Forecasting Financial Markets and Alternative Data Large-Scale Time-Series Data Engineering Census and Geographic Data Processing Pay: