OpenAI

Researcher, Safety Training, National Security

San Francisco · Washington, DC · US - Remote·$380K–500K

safety · AI

ABOUT THE TEAM

The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations.

We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely.

ABOUT THE ROLE

We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities.

IN THIS ROLE, YOU WILL

  • Research and implement methods for safety training, reinforcement learning, and adversarial robustness.
  • Develop evaluations, identify model failure modes, and use findings to improve training.
  • Work with research, engineering, security, and policy partners to support safe, reliable deployment.

YOU MIGHT THRIVE IN THIS ROLE IF YOU

  • Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness.
  • Have a degree in computer science, machine learning, or a related field
Read the rest on jobs.ashbyhq.com