Member of Technical Staff - Environments & Post-Training hillclimb San Francisco, CA Job Details Full-time 10 hours ago Benefits Health insurance Dental insurance Vision insurance 401(k) matching Food provided Qualifications Research Full Job Description About Hillclimb Hillclimb works with the world's best researchers, scientists, and engineers to teach AI to think like them so that it has the means to improve itself. We're a founding team of former DeepMind & quant researchers backed by Garry Tan, Jeff Dean, Tier 1 VCs & angel investors from OpenAI, Anthropic, DeepMind, SpaceXAI & Meta Superintelligence Labs. We achieved SOTA for math shortly after graduating YC.
About this role:
You're work will consist of the following and many more: Turn ML research problems into environments that reward a model for real progress and good research taste Test whether each environment's reward is repeatable and tracks the quality of the work it scores Build systems that create environments at scale and catch broken ones before they reach training Run post-training experiments to find which environments are trainable
Preferred Experience:
Experience with LLMs, RL, RLHF/RLAIF, RLVR, post-training, evals, graders, synthetic data, model training, or coding agents Understanding of what makes a reward evaluation trustworthy, and how models exploit one that is not Previous experience in designing automated systems for building, validating, and quality checking environments.
Benefits:
Extremely competitive salary and equity Health, Dental, Vision insurance Competitive 401k match Free lunch and dinner