AI researchers say safety is not keeping pace with new models
Current and former OpenAI and Google DeepMind researchers say labs keep releasing models while underestimating the time needed to assess increasingly autonomous systems. They describe internal pressure to move ahead and explain why some want development to slow.
Summary
Current and former OpenAI and Google DeepMind researchers say AI labs are developing models faster than they can prepare safeguards for increasingly autonomous systems. Their video testimonials were collected by the nonprofit Palisade Research and shared with Reuters.
Their warning goes beyond a distant prediction. Participants describe a tension inside labs: building more capable models is rewarded, while calls for more time to assess risks may carry less weight in decisions. They say their concerns are sincere, rather than a way to promote the technology.
In practice
The technical concern is recursive self-improvement. AI models already help researchers write code, run experiments, and analyse results. If a system eventually designs and develops a better model with little human involvement, that successor could accelerate the next cycle further. Researchers worry that the process could move faster than people can test, understand, and control each generation.
That scenario is different from what exists today. Anthropic says Claude wrote more than 80% of the code merged into its codebase in May 2026, but acknowledges that its systems cannot yet develop their own successors autonomously. The company also identifies persistent limits in deciding which research goals are worth pursuing.
The testimonials point to organisational obstacles as well. Rosie Campbell, a former OpenAI policy researcher, told Reuters that the company was becoming more divided into separate teams, making it harder for her to influence the technology’s direction before she left in 2024. Juan Felipe Ceron Uribe, an alignment researcher at OpenAI, describes competing labs pressing ahead despite uncertainty about the consequences.
What we still don't know
Participants do not necessarily agree on the size of the risk. Neel Nanda, a Google DeepMind researcher, gives his personal estimate in the videos that AI has at least a 10% chance of contributing to human extinction. That number is an individual judgment, not a probability established by scientific measurement.
Reuters also reports disagreement over what companies should do. Geoffrey Irving, a former OpenAI and DeepMind researcher, argues that a lab could slow down on its own. Industry leaders have emphasised the need for coordination among competitors. Anthropic’s Dario Amodei called this month for a slower pace in releasing new capabilities, and OpenAI’s Sam Altman publicly agreed; both companies have since released new models. OpenAI also said it had delayed the release of a more powerful model.
It has not been demonstrated that systems able to improve their successors autonomously will exist, or that the harms described will occur. The immediate question is who decides when a model has been assessed thoroughly enough for release, and what outside scrutiny can test that decision.
Why it matters
- The warning combines a technical possibility about future models with researchers’ firsthand accounts of decision-making inside labs.
- Using AI to accelerate AI research makes it harder to rely solely on companies’ internal timelines for safety assessment.
- Public calls to slow down can be judged against observable decisions: which models are released, which are delayed, and which evaluations are open to independent scrutiny.
