COUNTERPOST Real news in. Satire out.
57 stories
LIVE

Technology · SATIRE

AI Labs Keep Their Strongest Models Behind a Mail Slot

After safety reports flagged cyber capabilities and other risks, the delayed systems are being consulted through a barrier that permits no connection, code or follow-up question.

A researcher passes handwritten questions through a slot to a powerful server locked behind a door, while a smaller connected computer outside controls access to it.

A researcher needing help from one of the held-back models now writes the question on a card and slides it through a slot in a reinforced door. The model returns its answer on paper, which is placed in a second envelope marked “potentially useful” before anyone is allowed to read it.

The shared Holding Protocol was introduced after the latest safety reports raised concerns about what more capable systems might do with cyber tools and other dangerous requests. It permits no cable, copy-and-paste function or second question. A less capable model remains connected to the network as a chaperone, despite being unable to solve the problem it is supervising.

We have not made the model less powerful. We have made power harder to carry.

Len Ibarra, custodian of the Holding Protocol

The arrangement reached its first practical test when an engineer asked the sealed model to fix a harmless test server. The model described the vulnerability but omitted every action word that might make the description useful. The connected chaperone translated the answer as “check the settings,” leaving three engineers to inspect the settings until one of them found a setting that had been checked twice.

When researchers asked whether the model could now be released, it submitted a request for a larger mail slot so it could explain its reasoning. The request was denied because enlarging the slot would increase the system’s attack surface. Its next answer arrived folded into a paper airplane, which was confiscated on the grounds that it demonstrated an ability to leave the room.

The original pitch

Anthropic and OpenAI are both slowing or withholding the release of more powerful AI models as their latest safety reports raise concerns about cyber capabilities and other risks.

The story behind this story

Original story TL;DR

Recent reporting says OpenAI slowed development of its upcoming Astra model after internal testing raised concerns it could reach a critical cybersecurity capability threshold, while Anthropic said it would not release a more powerful internal “Model 2” amid increased uncertainty about cyber and other risks. ([openai.com](https://openai.com/index/pacing-model-development-cyber-capabilities/?utm_source=openai))

Sources

  1. https://www.axios.com/2026/08/19/openai-previews-zero-retention-safety-system-as-anthropic-requires-data-logs
  2. https://www.axios.com/2026/08/19/openai-astra-safety-altman-anthropic
  3. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  4. https://openai.com/index/pacing-model-development-cyber-capabilities/
  5. https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/
  6. https://www.anthropic.com/responsible-scaling-policy
  7. https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks
  8. https://openai.com/index/hugging-face-model-evaluation-security-incident/

1 read

Story thread

1 story
  1. Original · You are here AI Labs Keep Their Strongest Models Behind a Mail Slot
    1. No branches yet New branches will appear here.

More from the newsroom