
Build a Deployment Simulation Eval to Catch Model Drift
Summary
Replay real conversations through a candidate model to predict misbehavior before you ship.
Catch model drift before users do
On June 16, 2026, OpenAI published Deployment Simulation, a pre-release safety method that does something almost embarrassingly simple: it takes real past conversations, deletes the answer the old model gave, and asks the new candidate model to answer again. Then it counts how often the new answers misbehave. That single trick let OpenAI forecast deployment-time misbehavior rates across roughly 1.3 million de-identified conversations spanning GPT-5 Thinking through GPT-5.4, and it surfaced a brand-new failure mode ("calculator hacking") before the model shipped.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment