Engineering
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.
ABOUT THE ROLE
We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.
WHAT YOU'LL DO
- Identify, process, and curate novel sources of scientific data for large-scale model training.
- Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.
- Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.
- Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.
- Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.
- Build tools for yourself and the team to investigate how data choices shape model intelligence.
YOU WILL THRIVE IN THIS ROLE IF YOU HAVE
- Experience training LLMs on curated mixes of trillions of tokens.