Find Jobs
Find Jobs Near You – Available Work in Your Location
Applied Scientist III - VLM R&D
Career Insights for Chemist (General)
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on Washington data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
A Chemist studies the chemical and physical properties of substances or materials, focusing in one of many specialized areas. May work to develop new products or chemical processes.
$132,596 / year median in Washington
-0% projected decline
Job Description
The Opportunity We are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home. You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale. What You'll Do Follow the latest breakthroughs in multimodal and vision-language research from academia and industry, evaluate their relevance, and apply them to our in-house VLM development for smart home video understanding Train, fine-tune, and evaluate multimodal vision-language models on large-scale, real-world home video data Design and run rigorous evaluation pipelines to measure model quality on video understanding tasks such as event detection, activity recognition, and temporal reasoning Investigate user event patterns across tens of millions of authorized videos to inform model design and product direction Contribute to the architecture and training of a physical smart home foundation model, drawing on advances in visual transformers, physical world foundation models, and embodied AI Build rapid proofs of concept using AI-assisted research and development workflows, and carry promising directions from idea to validated prototype Publish research at top venues and contribute open-source models and datasets that help advance the community Help define research problems, set technical direction, and anticipate where academic research and industry solutions are heading What We're Looking For PhD in Computer Vision, Machine Learning, or a related field; or a Master's degree with a strong track record of research or applied impact (publications, open-source contributions, or shipped ML systems) Hands-on experience training and evaluating multimodal vision-language models Experience in one or more of: visual transformer algorithm innovation, physical world foundation models, or embodied AI Strong research sense: the ability to define the right problems, choose promising directions, and predict how research trends will translate into industry solutions Proficiency with AI-assisted research and fast POC development — you use modern AI tools to multiply your own research velocity Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), with the ability to work with large-scale video data pipelines Nice to Have Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar) Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment Experience deploying ML models in consumer products at scale