Data Scientist with Bachelor's degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor's degree in one of the aforementioned subjects.
Job Duties and Responsibilities:
Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to understand data, artificial intelligence, and analytics requirements. Contribute to translating business requirements into analytical tasks, technical specifications, success metrics, acceptance criteria, and implementation steps. Collect, clean, profile, analyze, and interpret structured, semi-structured, and unstructured datasets to identify trends, patterns, anomalies, and business opportunities. Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering using established analytical methods. Develop, train, tune, test, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, and anomaly-detection use cases. Compare model performance using metrics such as precision, recall, F1 score, AUC, RMSE, and MAE, and prepare dashboards, reports, and visualizations to communicate the results. Develop and support Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models. Apply prompt-engineering, grounding, citation-validation, structured-output, and response-evaluation techniques to improve the accuracy and reliability of Generative AI applications. Develop, test, and support single-agent and multi-agent workflows using LangChain, LangGraph, LlamaIndex, or comparable Agentic AI frameworks. Support human-in-the-loop approvals, confidence thresholds, fallback mechanisms, retry policies, and testing for hallucinations, bias, prompt injection, sensitive-data exposure, and unauthorized tool usage. Develop Python-based components and reusable APIs for data processing, machine learning, Generative AI, and Agentic AI applications. Contribute to the development and maintenance of data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, and data lakes. Develop and support ETL and ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and analytical preparation. Support batch, streaming, Change Data Capture, and event-driven data pipelines, including data-quality checks, schema validation, logging, metadata management, and monitoring. Participate in testing, deployment, versioning, monitoring, troubleshooting, and documentation of machine-learning models, AI applications, and data pipelines while coordinating with cross-functional, onshore, and offshore teams. Technologies Involved / Skills required for the position: Effective communication, collaboration, analytical, problem-solving, and documentation skills for working with cross-functional technical and business teams. Ability to understand business requirements and translate them into analytical tasks, technical specifications, success metrics, acceptance criteria, and implementation steps. Proficiency in Python and SQL, with experience using Pandas, NumPy, PySpark, or comparable technologies to collect, clean, profile, transform, and analyze structured, semi-structured, and unstructured data. Working knowledge of exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature-engineering techniques. Experience developing and evaluating regression, classification, clustering, forecasting, recommendation, and anomaly-detection models using Scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch, or comparable frameworks. Knowledge of model-evaluation metrics, including precision, recall, F1 score, AUC, RMSE, and MAE, with experience creating dashboards, reports, and visualizations using Power BI, Tableau, Looker, Omni, or comparable tools. Experience developing Retrieval-Augmented Generation solutions using embeddings, semantic search, hybrid search, reranking, vector databases, large language models, and enterprise data. Working knowledge of prompt engineering, grounding, citation validation, structured outputs, response evaluation, and guardrails for Generative AI applications. Experience developing and testing single-agent and multi-agent workflows using LangChain, LangGraph, LlamaIndex, or comparable Agentic AI frameworks. Knowledge of human-in-the-loop processes, confidence thresholds, fallback mechanisms, retry policies, hallucination evaluation, bias detection, prompt-injection prevention, and sensitive-data protection. Proficiency in developing Python-based data-processing, machine-learning, Generative AI, and Agentic AI components, with knowledge of reusable APIs using FastAPI, Flask, or comparable frameworks. Experience working with data models and analytical schemas using MongoDB, PostgreSQL, MySQL, SQL Server, BigQuery, Snowflake, Databricks, or comparable database and cloud data technologies. Experience developing and supporting ETL and ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and analytical preparation. Working knowledge of batch, streaming, Change Data Capture, and event-driven pipelines using technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, or EventBridge. Understanding of testing, deployment, versioning, monitoring, troubleshooting, and documentation practices for machine-learning models, AI applications, and data pipelines, including familiarity with Docker, CI/CD, MLOps, and LLMOps. Work location is Portland, ME with required travel to client locations throughout USA. Rite Pros is an equal opportunity employer (EOE).
Please Mail Resumes to:
Rite Pros, Inc. 565 Congress St, Suite # 305 Portland, ME 04101.