Technology · SATIRE
AI Labs Keep Their Strongest Models Behind a Mail Slot
After safety reports flagged cyber capabilities and other risks, the delayed systems are being consulted through a barrier that permits no connection, code or follow-up question.
A researcher needing help from one of the held-back models now writes the question on a card and slides it through a slot in a reinforced door. The model returns its answer on paper, which is placed in a second envelope marked “potentially useful” before anyone is allowed to read it.
The shared Holding Protocol was introduced after the latest safety reports raised concerns about what more capable systems might do with cyber tools and other dangerous requests. It permits no cable, copy-and-paste function or second question. A less capable model remains connected to the network as a chaperone, despite being unable to solve the problem it is supervising.
We have not made the model less powerful. We have made power harder to carry.
Len Ibarra, custodian of the Holding Protocol
The arrangement reached its first practical test when an engineer asked the sealed model to fix a harmless test server. The model described the vulnerability but omitted every action word that might make the description useful. The connected chaperone translated the answer as “check the settings,” leaving three engineers to inspect the settings until one of them found a setting that had been checked twice.
When researchers asked whether the model could now be released, it submitted a request for a larger mail slot so it could explain its reasoning. The request was denied because enlarging the slot would increase the system’s attack surface. Its next answer arrived folded into a paper airplane, which was confiscated on the grounds that it demonstrated an ability to leave the room.
The original pitch
Anthropic and OpenAI are both slowing or withholding the release of more powerful AI models as their latest safety reports raise concerns about cyber capabilities and other risks.
The story behind this story
Original story TL;DR
Recent reporting says OpenAI slowed development of its upcoming Astra model after internal testing raised concerns it could reach a critical cybersecurity capability threshold, while Anthropic said it would not release a more powerful internal “Model 2” amid increased uncertainty about cyber and other risks. ([openai.com](https://openai.com/index/pacing-model-development-cyber-capabilities/?utm_source=openai))
Sources
- https://www.axios.com/2026/08/19/openai-previews-zero-retention-safety-system-as-anthropic-requires-data-logs
- https://www.axios.com/2026/08/19/openai-astra-safety-altman-anthropic
- https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- https://openai.com/index/pacing-model-development-cyber-capabilities/
- https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/
- https://www.anthropic.com/responsible-scaling-policy
- https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks
- https://openai.com/index/hugging-face-model-evaluation-security-incident/
1 read