Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Collabera LLC

Data Engineer

Career Insights for Data Engineer

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on national data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Data Engineer designs, builds and manages the information or big data infrastructure. Develops the architecture that helps analyze and process data in the way the organization needs it. Makes sure those systems are performing smoothly.

$111,705 / year median in the U.S.

+21% projected growth

Explore Career

Job Description

Job Description

Data Engineer Remote:

Houston ,

Texas, US Salary Range:

135000.00

•165000.00 |

Per Annum Job Code:

371658

End Date:

2026-10-15

Days Left:

28 days, 7 hours left

Title:

Senior Data Engineer

•AI & Platform

Location:

US

•Remote (EST/CST)

Employment:

Full-Time opportunity

Salary Range:

$135k/ann

•$165k/ann.

Key Responsibilities:

Design, build, and maintain scalable, self-healing data pipelines using Databricks, Python, PySpark, and SQL.

Develop data pipelines across Bronze, Silver, and Gold/Certified Gold layers, ensuring data quality and reliability.

Ingest, transform, and serve data from dozens of source systems, including enterprise applications, financial systems, IoT, web/mobile analytics, and third-party platforms.

Build error recovery and quarantine workflows to isolate failed records while allowing valid data to continue through the pipeline.

Implement schema validation, data quality checks, anomaly detection, automated testing, regression testing, and data observability.

Develop and maintain data models with proper data grain, keys, referential integrity, lineage, and governance.

Build infrastructure that supports AI/ML workloads, including feature stores, embedding pipelines, vector search, and real-time serving layers.

Build and maintain MCP server integrations that expose enterprise data to LLM-powered tools and AI agents.

Develop integrations that allow AI applications to connect to, query, and retrieve data through MCP servers.

Design and implement RAG (Retrieval-Augmented Generation) architectures and integrate LLM-powered solutions with enterprise data.

Work with vector databases such as Pinecone, Weaviate, or similar technologies.

Support AI model training, evaluation, deployment, monitoring, and productionization in partnership with Data Science and Product teams.

Evaluate AI-powered data engineering and data quality tools, including solutions for automated schema detection, cataloging, completeness checks, and data validation.

Use AI-assisted testing and evaluation approaches to validate data and AI outputs against business requirements.

Develop APIs and data interfaces that enable AI products and internal applications to query and interact with data in real time.

Implement data governance practices covering access controls, PII handling, data classification, compliance, and appropriate AI data usage.

Build monitoring, alerting, SLA tracking, and data freshness capabilities into data platforms.

Document data models, pipeline architectures, AI integrations, and reusable engineering patterns.

Required Qualifications:
  • 5+ years of experience in Data Engineering, working with data from multiple enterprise source systems.
  • Strong hands-on experience with Databricks.
  • Strong Python and PySpark development experience.
  • Strong SQL skills.
  • 2-3+ years of hands-on experience with MCP Servers / Model Context Protocol.
  • Hands-on experience building and implementing RAG models/architectures.
  • Experience connecting to, querying, and integrating data through MCP servers.
  • Strong experience building self-healing or fault-tolerant data pipelines and error recovery workflows.
  • Experience with schema validation and data quality frameworks.
  • Strong understanding of Bronze/Silver/Gold data architecture.
  • Experience with LLM integration patterns, AI agents, tool-use frameworks, and AI-enabled data solutions.
  • Experience with vector databases such as Pinecone, Weaviate, or equivalent.
  • Experience building data pipelines and infrastructure suitable for AI/ML workloads.

Understanding of data governance, lineage, monitoring, observability, and data quality. This is a direct hire opportunity. The selected candidate will be employed directly by our client. All compensation and benefits, including but not limited to medical insurance, retirement plans, paid time off, and other perks, will be provided by the client in accordance with their internal policies and subject to applicable laws and eligibility requirements. Job Requirement

MCP

RAG

Pyspark

Reach Out to a Recruiter

Recruiter

Email

Phone

Christin Mathew

christin.mathew@collabera.com

Benefits

  • Paid Time Off (PTO)
  • Dental Insurance
  • Other Retirement and Savings