Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

Two Six Technologies

Data Collection Engineer

Career Insights for Data Entry Clerk

See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.

Scorecard

Based on national data

Review key factors to help you decide if this role fits your goals. How is this calculated?

Were these scores useful?

What they do

A Data Entry Clerk types or scans information into a computer to create databases, programs or summary data reports. Transcribes handwritten information from order forms, receipts or other documents to type into a computerized record, or enters data for medical records, or enters data in code to update company websites.

$42,814 / year median in the U.S.

-35% projected decline

Explore Career

Job Description

At Two Six Technologies, we build, deploy, and implement innovative products that solve the world's most complex challenges today. Through unrivaled collaboration and unwavering trust, we push the boundaries of what's possible to empower our team and support our customers in building a safer global future. Two Six Technologies is seeking a highly skilled Data Collection Engineer to design, scale, and maintain our distributed web scraping and data extraction infrastructure. In this role, you will be responsible for building resilient data pipelines that harvest data from complex web ecosystems, ensuring strict data quality through automated validation, and managing containerized workloads at scale. If you thrive on reverse-engineering web applications, overcoming anti-bot barriers, and orchestrating distributed systems, we want you on our team.
Location:
100% Remote What you will
Do:
Distributed Crawler Development:
Design and deploy high-performance, distributed web scrapers using Python and Scrapy to extract massive datasets efficiently.
Dynamic Content Extraction:
Utilize Browser Scripting tools to navigate, interact with, and extract data from modern, dynamic, and JavaScript-heavy websites.
Infrastructure & Container Orchestration:
Deploy, scale, and manage scraping workloads on Kubernetes , ensuring optimal resource allocation and fault tolerance.
Data Validation & Quality Assurance:
Define strict JSON Schemas and leverage Pydantic to enforce data types, validate incoming payloads, and catch data drift early.
Data Ingestion & Storage:
Build and optimize search and storage pipelines using Elasticsearch , transforming raw web dumps into highly structured, searchable data.
Pipeline Workflow Management:
Architect robust pipeline workflows to manage the end-to-end data lifecycle—from discovery and extraction to validation and storage.
Anti-Bot & Proxy Engineering:
Manage complex proxy rotation, session handling, and browser fingerprinting to maintain high success rates against advanced anti-scraping systems. What you will need (basic qualifications):
Experience:
7+ years of professional software engineering experience, with a heavy focus on web scraping, data engineering, or distributed systems.
Analytical Mindset:
Excellent reverse-engineering skills, with the ability to dissect network traffic, unearth hidden APIs, and bypass complex web barriers.
Reliability Focus:
A strong commitment to data integrity, system monitoring, and building self-healing scraping systems.
Core Language:
Expert-level proficiency in Python .
Scraping Frameworks:
Deep experience with Scrapy and distributed scraping architectures (e.g., handling distributed queues, broad vs. deep crawling).
Automation & Browser Scripting:
Proven experience with browser automation tools (Playwright, Selenium, or Puppeteer).
Data Serialization & Validation:
Mastery of
JSON , JSON
Schema , and data validation using Pydantic .
Search & Analytics Engines:
Hands-on experience indexing, querying, and optimizing Elasticsearch clusters.
Orchestration:
Strong proficiency in managing and scaling applications within Kubernetes environments.
Workflow Management:
Experience building structured pipeline workflows to handle complex, multi-stage data extraction tasks.
Education:
Bachelor's degree in Computer Science, Engineering Nice if you
Have:
AI & Intelligent Extraction:
Experience leveraging LLMs or Computer Vision for adaptive scraping, parsing unstructured HTML, or bypassing CAPTCHAs ( AI in data collection ).
Cloud Infrastructure:
Strong hands-on experience with AWS ecosystems (e.g., EKS, EC2, S3, RDS).
Relational Databases:
Proficiency in SQL for querying, schema design, and storing structured relational data.
In-Memory Data Structures:
Experience with Redis (specifically for caching, deduplication, or as a Scrapy distributed queue back-end).
Event Streaming:
Familiarity with Apache Kafka for real-time data streaming and decoupled pipeline architectures.
Containerization:
Strong foundation in Docker for local development and containerizing scraping microservices.
DevOps:
Experience with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins) for automated testing and deployment of crawlers.
Clearance Requirement:
Eligible to obtain a clearance Looking for other great opportunities? Check out Two Six Technologies Opportunities for all our Company's current openings! Ready to make the first move towards growing your career? If so, check out the Two Six Technologies Candidate Journey! This will give you step-by-step directions on applying, what to expect during the application process, information about our rich benefits and perks along with our most frequently asked questions. If you are undecided and would like to learn more about us and how we are contributing to essential missions, check out our Two Six Technologies News page! We share information about the tech world around us and how we are making an impact! Still have questions, no worries! You can reach us at Contact Two Six Technologies. We are happy to connect and cover the information needed to assist you in reaching your next career milestone. Two Six Technologies is an Equal Opportunity Employer and does not discriminate in employment opportunities or practices based on race (including traits historically associated with race, such as hair texture, hair type and protective hair styles (e.g., braids, twists, locs and twists)), color, religion, national origin, sex (including pregnancy, childbirth or related medical conditions and lactation), sexual orientation, gender identity or expression, age (40 and over), marital status, disability, genetic information, and protected veteran status or any other characteristic protected by applicable federal, state, or local law. For more information review the Two Six Technologies Equal Employment Opportunity and Affirmative Action Policy and the EEO Poster. If you are an individual with a disability and would like to request reasonable workplace accommodation for any part of our employment process, please send an email to accommodations@twosixtech.com. Information provided will be kept confidential and used only to the extent required to provide needed reasonable accommodations. Additionally, please be advised that this business uses E-Verify in its hiring practices. By submitting the following application, I hereby certify that to the best of my knowledge, the information provided is true and accurate.