Research · Pre-Training Data
OUR MISSION
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
About the Role
- Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages.
- Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
- Design and run scientific experiments to advance our understanding of scaling large language models to improve multilingual data efficiency.
- Lead small research projects independently while collaborating on larger initiatives.
- Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain, and close those gaps with targeted data.
- Work with pre-training, mid-training, and post-training teams to land measurable, step-function improvements in multilingual capability.
About You
- Strong software engineering fundamentals and comfort processing web-scale datasets in distributed environments.
- Experience building large-scale data pipelines for language models, machine translation, speech, or search — ideally covering more than one language.
- Fluency or working proficiency in at least one language other than English, and genuine curiosity about how languages differ.
- A rigorou