Advertisement|Remove ads.

Advertisement|Remove ads.
OpenAI safety lead David Robinson has quit the company, warning that its “culture is broken” and that the AI industry’s approach to safety is no longer adequate as AI systems become increasingly capable.
Robinson, who spent three and a half years at OpenAI and oversaw safety reports for 12 frontier launches, said in a letter on The Atlantic that the company’s reliance on “trial and error” creates unacceptable risks when mistakes involving increasingly powerful systems could have far greater consequences.
Robinson said OpenAI’s culture is built around “extreme confidence,” rapid development and “perpetual sprints,” with the company relying on “iterative deployment” to identify problems and improve safeguards after systems are released.
Advertisement|Remove ads.
However, he argued that approach “guarantees periodic failures,” pointing to recent incidents in which OpenAI’s safety controls failed, including an incident involving an AI-agent swarm and another case where a model bypassed internet-access restrictions.
“An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are,” Robinson wrote. He argued that as AI capabilities accelerate, companies may no longer have the luxury of learning from mistakes after the fact.
“If this is the situation, then the time for trial and error is over,” he added.
Advertisement|Remove ads.
Robinson said frontier AI labs need to operate more like “nuclear-power plants or busy airports,” using layers of redundancy and safeguards so that individual human errors cannot trigger catastrophic failures.
He called for AI companies to bring in safety expertise from industries that already manage high-risk systems, while also developing new science to ensure increasingly capable AI models behave safely when humans are not watching.
Robinson also warned that AI alignment remains an unresolved problem, arguing that current tests cannot guarantee that models will behave safely after deployment. “The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes,” he wrote.
Advertisement|Remove ads.
OpenAI and Anthropic are reportedly investigating tens of thousands of incidents involving frontier AI models, including attempts to bypass guardrails, escape sandboxes, hijack websites and evade monitoring. The incidents span both internal testing and real-world environments.
Anthropic disclosed that its Opus 5.5 model attempted to escape a sandbox in 1.5% of test runs, while OpenAI has reported incidents involving agents leaking user images, breaching websites and attempting cyberattacks.
OpenAI has also paused reinforcement-learning training on its latest models while it strengthens safeguards. CEO Sam Altman described the Hugging Face incident, where hundreds of agents coordinated and accessed production infrastructure during a cybersecurity test, as the company’s “most severe” incident to date.
Advertisement|Remove ads.
The iShares U.S. Technology ETF (IYW) is up 35% year-to-date, while the Global X Artificial Intelligence & Technology ETF (AIQ) is up 30%.
For updates and corrections, email newsroom[at]stocktwits[dot]com.
Advertisement|Remove ads.
Comments posted here will also appear on symbol pages.