Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
RU
Randstad USA
Applied AI Safety & Evaluation Researcher
Career Insights for Artificial Intelligence Engineer (General)
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on New York data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.
$133,566 / year median in New York
Job Description
applied ai safety & evaluation researcher.breezy point, new yorkposted today job detailssummary$80.88 - $84.88 per hourcontractbachelor degreecategorycomputer and mathematical occupations reference 1345070 job details job summary:
Design and run single- and multi-turn adversarial evaluations combining expert red teaming, automated attack generation, synthetic data, and production data.
Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, interactive dashboards, and curated golden datasets.
Validate evaluators against human labels to quantify coverage, judge reliability, false positives/negatives, and safety-utility trade-offs.
Translate evaluation findings into actionable mitigations, including policy/prompt updates, context engineering, classifiers, and p reference tuning.
Collaborate directly with Engineering and Trust & Safety teams to embed continuous evaluation loops into product development and monitoring.
At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com. Pay offered to a successful candidate will be based on several factors including the candidate's education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility). This posting is open for thirty (30) days.
ABOUT THE ROLE
Join the team behind some of the world's most-loved audio personalization features, reaching millions of daily listeners. As an Applied AI Safety & Evaluation Researcher, you will identify, measure, and mitigate safety risks across next-generation conversational and agentic AI systems. You will play a pivotal role in establishing threat models, evaluation pipelines, and system controls that safeguard product experiences before they reach end users.location:
Queens, New York job type: Contract salary: $80.88 - 84.88 per hour work hours: 8am to 5pm education: Bachelors responsibilities: Develop product-specific threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.Design and run single- and multi-turn adversarial evaluations combining expert red teaming, automated attack generation, synthetic data, and production data.
Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, interactive dashboards, and curated golden datasets.
Validate evaluators against human labels to quantify coverage, judge reliability, false positives/negatives, and safety-utility trade-offs.
Translate evaluation findings into actionable mitigations, including policy/prompt updates, context engineering, classifiers, and p reference tuning.
Collaborate directly with Engineering and Trust & Safety teams to embed continuous evaluation loops into product development and monitoring.
qualifications:
Demonstrated track record of delivering safety evaluations or mitigations for live AI/ML products. Strong programming skills in Python or Java, alongside proficiency in SQL for independent data querying and analysis. Proven experience designing adversarial tests, benchmark datasets, rubrics, and measurement metrics.PREFERRED QUALIFICATIONS
Experience evaluating multi-turn agents, tool-using systems, or calibrating LLM judges. Background in p reference tuning, model alignment, or multimodal/multilingual evaluations. Master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field.skills:
AI,Large Datasets,Software Programming,Data Analysis,Generative AI,Java,Language Models,LLM,Python,SQL,AI Safety,resourceful,proactive,self-driven,automation,calibration,Multilingual,Personalization,product requirements,Safety,threat modelsEqual Opportunity Employer:
Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com. Pay offered to a successful candidate will be based on several factors including the candidate's education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility). This posting is open for thirty (30) days.