OpenAI
Predicting LLM Safety Before Release by Simulating Deployment
This paper introduces deployment simulation, a method for predicting LLM safety before release by resampling the next assistant response from de-identified production conversation prefixes using a candidate model. The authors evaluate this approach across GPT-5-series deployments, finding that it produces informative estimates of post-deployment…
Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, et al.- Published
- Jul 2026
- Citations
- 1
- Code
- Not linked
