Safeguards Enforcement Analyst, Conventional Weapons
San Francisco, CA | New York City, NY | Washington, DC
Safeguards (Trust & Safety)
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
As a Safeguards Enforcement Analyst focused on Conventional Weapons, your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, specifically utilizing conventional weapons and dangerous technology. You will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas.
Important context for this role: In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature.
Key responsibilities
Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically so