Find Jobs
Find Jobs Near You – Available Work in Your Location
Data Engineer
Career Insights for Data Engineer
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on national data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Data Engineer designs, builds and manages the information or big data infrastructure. Develops the architecture that helps analyze and process data in the way the organization needs it. Makes sure those systems are performing smoothly.
$111,705 / year median in the U.S.
+21% projected growth
Job Description
Job Description
Data Engineer Remote:
Houston ,
Texas, US Salary Range:
135000.00
•165000.00 |
Per Annum Job Code:
371658
End Date:
2026-10-15
Days Left:
28 days, 7 hours left
Title:
Senior Data Engineer
•AI & Platform
Location:
US
•Remote (EST/CST)
Employment:
Full-Time opportunity
Salary Range:
$135k/ann
•$165k/ann.
Key Responsibilities:
Design, build, and maintain scalable, self-healing data pipelines using Databricks, Python, PySpark, and SQL.
Develop data pipelines across Bronze, Silver, and Gold/Certified Gold layers, ensuring data quality and reliability.
Ingest, transform, and serve data from dozens of source systems, including enterprise applications, financial systems, IoT, web/mobile analytics, and third-party platforms.
Build error recovery and quarantine workflows to isolate failed records while allowing valid data to continue through the pipeline.
Implement schema validation, data quality checks, anomaly detection, automated testing, regression testing, and data observability.
Develop and maintain data models with proper data grain, keys, referential integrity, lineage, and governance.
Build infrastructure that supports AI/ML workloads, including feature stores, embedding pipelines, vector search, and real-time serving layers.
Build and maintain MCP server integrations that expose enterprise data to LLM-powered tools and AI agents.
Develop integrations that allow AI applications to connect to, query, and retrieve data through MCP servers.
Design and implement RAG (Retrieval-Augmented Generation) architectures and integrate LLM-powered solutions with enterprise data.
Work with vector databases such as Pinecone, Weaviate, or similar technologies.
Support AI model training, evaluation, deployment, monitoring, and productionization in partnership with Data Science and Product teams.
Evaluate AI-powered data engineering and data quality tools, including solutions for automated schema detection, cataloging, completeness checks, and data validation.
Use AI-assisted testing and evaluation approaches to validate data and AI outputs against business requirements.
Develop APIs and data interfaces that enable AI products and internal applications to query and interact with data in real time.
Implement data governance practices covering access controls, PII handling, data classification, compliance, and appropriate AI data usage.
Build monitoring, alerting, SLA tracking, and data freshness capabilities into data platforms.
Document data models, pipeline architectures, AI integrations, and reusable engineering patterns.
Required Qualifications:
- 5+ years of experience in Data Engineering, working with data from multiple enterprise source systems.
- Strong hands-on experience with Databricks.
- Strong Python and PySpark development experience.
- Strong SQL skills.
- 2-3+ years of hands-on experience with MCP Servers / Model Context Protocol.
- Hands-on experience building and implementing RAG models/architectures.
- Experience connecting to, querying, and integrating data through MCP servers.
- Strong experience building self-healing or fault-tolerant data pipelines and error recovery workflows.
- Experience with schema validation and data quality frameworks.
- Strong understanding of Bronze/Silver/Gold data architecture.
- Experience with LLM integration patterns, AI agents, tool-use frameworks, and AI-enabled data solutions.
- Experience with vector databases such as Pinecone, Weaviate, or equivalent.
- Experience building data pipelines and infrastructure suitable for AI/ML workloads.
Understanding of data governance, lineage, monitoring, observability, and data quality. This is a direct hire opportunity. The selected candidate will be employed directly by our client. All compensation and benefits, including but not limited to medical insurance, retirement plans, paid time off, and other perks, will be provided by the client in accordance with their internal policies and subject to applicable laws and eligibility requirements. Job Requirement
MCP
RAG
Pyspark
Reach Out to a Recruiter
Recruiter
Phone
Christin Mathew
christin.mathew@collabera.com
Benefits
- Paid Time Off (PTO)
- Dental Insurance
- Other Retirement and Savings