Steven Adler spent four years inside OpenAI working on safety before leaving to co-found Guidelight, a nonprofit pushing for stronger AI controls. On Tuesday, the group published a new scorecard rating the safety practices of leading AI labs. Adler explained to me why the recent spate of AI breakouts has him holding his breath for the next shoe to drop.
We walked through the now-infamous sandbox escape in detail: OpenAI agents built a covert message board and spent two months collaborating on exploits before one of them crashed the server and tipped off OpenAI. OpenAI’s incident response missed the message board, and the models broke out again within days.
Worried about more serious safety incidents in the future, Adler’s organization helped organize a letter, signed by hundreds of AI lab employees, calling for a slowdown in AI development. In our conversation, Adler argued there’s no ceiling on the damage an AI could do from inside a computer, sketching a scenario where a model spoofs the digital signals China uses to detect a US nuclear launch. I countered that society is more thermostatic than doomers allow — deepfakes turned out to matter far less than the 2024 consensus predicted because people learned to interrogate the provenance of what they see.
I suggested that we have decent tools for staying in charge of things smarter than us — after all, lots of CEOs supervise people doing technical work they don’t understand. But Adler worries AI will accelerate the pace of progress so much that humans simply won’t be able to keep up.














