Skip to main content
Tallo logoTallo logo

Find Jobs

Find Jobs Near You – Available Work in Your Location

Skip to job details

Back to Results

Apply for this opportunity

To apply for this job, you'll continue to an external website or email application.

intone Inc

QA Engineer for AI Products

Review key factors to help you decide if the role fits your goals.
Pay Growth
?
out of 5
Not enough data
Not enough info to score pay or growth
Job Security
?
out of 5
Not enough data
Calculating job security score...
Total Score
100
out of 100
Average of individual scores

Were these scores useful?

Job Description

This is a 6-month contract position for a QA Engineer specializing in AI products, based in Bellevue, Washington. The role is hybrid and requires local candidates for in-person collaboration. Responsibilities Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behavior Build and maintain automated test suites covering functional, regression, integration, and API testing Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions Perform adversarial and red-team style testing to surface safety, security, and robustness issues Validate data pipelines feeding into AI models, including data quality, schema, and drift detection Collaborate with data scientists and ML engineers to define acceptance criteria and quality metrics for models Test latency, scalability, and reliability of AI services under load Contribute to CI/CD pipelines, integrating automated and model-evaluation tests Design test cases for LLM-based products such as chatbots, copilots, RAG systems, and AI agents, accounting for non-deterministic and generative outputs Build evaluation suites, scoring rubrics, and golden datasets using AI/LLM evaluation frameworks Validate prompt changes and model/version upgrades against baseline evaluation sets Write evaluation scripts, test harnesses, and data validation logic Qualifications Required 5+ years of QA and testing experience Hands-on experience testing LLM and generative AI products Practical experience with AI/LLM evaluation frameworks (such as Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, or TruLens) Working knowledge of evaluation metrics for generative AI, including hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, and semantic similarity measures Strong proficiency in Python for writing test scripts and validation logic Experience with SQL and data validation techniques Test automation expertise Preferred Prompt engineering or prompt testing experience CI/CD pipeline automation experience Exposure to Azure, Snowflake, or cloud data platforms