The government built a vault around something that already left the building. Here’s the engineering proof.

This is Part 4 — the finale. Part 1 by Eric Mitchell: the political case. Part 2: the engineering case — the framework has no test suite. Part 3: the financial tell — why Altman’s offer is the survival math of a founder who knows the bubble knows it’s a bubble. This piece closes the series with the irony sitting under all of it: they may win every institutional fight and still lose, because the thing they’re trying to contain was never containable in the first place.

“The government built the most consequential gate in the history of American technology policy. It controls what models ship, to whom, on what timeline. It can trigger a global recall. It can restore partial access through a single letter. And it has no test suite.”

— Part 2 of this series. The gate has no criteria. This piece goes one level deeper: even with criteria, the gate can’t hold. Here’s why.


 The Premise of Containment Has Already Failed

Part 3 ended with a question nobody in Washington will say plainly: what happens when the government wins the institutional fight — the off-switch, the equity stake, the regulatory framework — and the thing it was fighting to contain has already left the room?

That’s not a hypothetical. It’s the current state of the technology.

The government’s model of AI containment assumes that capability lives in a controlled artifact — a specific model, deployed by a specific company, accessible through a specific API — and that controlling access to that artifact controls the capability. Recall the model, control the capability. Gate the deployment, control the capability. Take a stake in the company, align the incentives.

Every piece of this framework is wrong. Not wrong in implementation — wrong in premise. Capability in frontier AI doesn’t live in a controlled artifact. It propagates through the ecosystem through a mechanism the government’s own researchers have documented in precise detail, and which no proposed regulatory measure fully addresses. The mechanism is called distillation. And it means the vault was empty before the lock was installed.


 What Distillation Actually Does

Distillation is how you take capability from a large, expensive model and transfer it to a smaller, cheaper one — not by copying the weights, but by using the large model’s outputs to train the smaller model. You prompt the capable model, collect its responses, and use those responses as training data. The student model learns to imitate the teacher’s behavior without ever touching the teacher’s weights. [8].

This is how most of the open-source AI ecosystem develops. It’s how Meta’s Llama models have been refined, how DeepSeek built competitive reasoning capability at a fraction of U.S. compute costs, and how virtually every efficient frontier-competitive model in 2026 was developed.[10;13].

It’s also, as the Center for a New American Security documented in a major policy report this year, how China’s AI ecosystem has been systematically extracting capability from U.S. frontier models at industrial scale.

Here is the part that breaks the containment premise:

Distillation doesn’t require access to the weights. It requires access to the outputs.

Anthropic, Google, and OpenAI have documented that named Chinese entities — DeepSeek, Moonshot, MiniMax — together generated over 16 million exchanges with U.S. models, representing an estimated 150 to 400 billion tokens of extracted capability [5].

DeepSeek-R1’s entire supervised fine-tuning dataset is estimated at 6.4 billion tokens.

The adversarial campaigns extracted more capability than that model’s entire training dataset — not by stealing anything, but by asking questions through APIs that were openly available.

The government recalled Anthropic’s models and demanded zero jailbreaks before restoration. Meanwhile, the capability those models represent had been flowing out through commercial APIs for months, in a form that requires no jailbreak at all — just a subscription and a well-structured prompt.


 The Front Door Was Always Open

The CNAS report on adversarial distillation describes the infrastructure in detail that should end any serious policy conversation about containment through model access restriction. The adversarial distillation supply chain runs through commercial token mixers — services like OpenRouter that aggregate access to multiple models through a single API endpoint.

It runs through “hydra cluster” architectures: distributed networks of fraudulent accounts where any single disabled account is immediately replaced by another. One proxy network operated more than 20,000 fraudulent accounts in parallel. When Anthropic released a new model during an active campaign, the Chinese entity conducting it pivoted within 24 hours, redirecting nearly half its traffic to capture capabilities from the latest version.

This is not hacking. It is not a cyberattack. It is a sophisticated use of commercially available services. The “vault” the government is building around frontier AI models is a vault whose contents are sold at the front counter.

The proposed Deterring American AI Model Theft Act unanimously cleared the House Foreign Affairs Committee in April 2026 + [5].

NSTM-4, issued by the White House’s own Office of Science and Technology Policy the same month, found that “foreign entities, principally based in China, are engaged in deliberate, industrial-scale campaigns to distill U.S. frontier AI systems.” [5] The government knows this is happening. It is building a gate that restricts access for American developers, researchers, and allied nations while doing essentially nothing to stop the adversarial extraction that motivated the framework in the first place.

The Fable 5 recall constrained more than 100 American institutions for fourteen days.
The Chinese distillation campaigns it was supposedly designed to address continued operating through token mixers the entire time. [6;15;11;10]


 Weights Are Just Bits, and Bits Copy

There is a second, harder version of this problem that doesn’t require adversarial actors at all.

Model weights are files. They are very large files — the weights for a frontier-scale model can run to hundreds of gigabytes — but they are files. They can be copied, stored, transferred, and deployed by anyone with the hardware to run them. Once a capability is trained into a set of weights and those weights exist anywhere outside an air-gapped facility, containment is a matter of degree, not kind.

The open-weight frontier in 2026 makes this concrete. DeepSeek V4-Pro, with 1.6 trillion parameters and 49 billion active, is the largest open-weight model available — and it is publicly downloadable.

Llama 4, Qwen 3.6, and multiple other models with frontier-competitive reasoning capability are openly available and can be run locally, fine-tuned without restriction, and deployed without any API that a government could monitor or gate. [11]. The gap between open-weight and proprietary capability has closed to the point where leading technical analysts project open-weight models will match proprietary alternatives on the majority of practical tasks by late 2026.

A regulatory framework built on controlling access to specific proprietary models is a framework that becomes strategically irrelevant as open-weight alternatives reach parity. The government is building a gate in front of one door in a building with no walls.


 What My Published Architecture Already Says About This

I want to be direct about the timeline, because it matters.

On April 2026 — seven weeks before the Fable 5 recall, two months before EO 14409 — I published a piece on this Substack arguing that the story of Mythos finding a 17-year-old FreeBSD exploit wasn’t about the vulnerability. It was about what the finding revealed: that capability in these systems emerges from training in ways that are not fully predictable, cannot be designed out, and cannot be removed after the fact without destroying the system’s usefulness.

The constraint layer sits downstream of the capability structure. Safety training intercepts outputs. It doesn’t modify weights. The geometry runs to completion; the filter redirects at the end. A government framework that treats a jailbreak as a patchable vulnerability is misunderstanding the architecture at the level that determines whether any of its actions have any effect [15].

The AI Workflow Architect Worksheet I published in March says that any gate without defined advancement criteria, collapse conditions, and recovery moves isn’t a gate — it’s a vibes-based sequence with a deploy button on the end. Part 2 showed that EO 14409 fails that standard on every row.

This piece adds the third floor: even a gate that passed that standard would be gating a capability that is already distributed through distillation, already encoded in open-weight models, and already operating in adversarial hands. A perfect gate on an empty vault is still an empty vault.


The Institutional Win, The Strategic Loss

Eric’s Part 3 read the Altman offer correctly: it’s the survival math of someone who watched Washington demonstrate an off-switch and decided a government partner was cheaper than a government adversary.[cite:236] The bubble knows it’s a bubble. The offer is insurance, not generosity.

Run the institutional logic forward. The government gets an equity stake. It’s on the cap table. It collects the dividend. The regulator and the regulated are fused at the balance sheet. By every measure of Washington’s stated objectives — American AI companies under American oversight, sensitive capability within the U.S. regulatory perimeter, frontier AI development controlled by an accountable party — this is a win.

Now ask what that win actually controls.

It controls the API. It controls the deployment pipeline. It controls what a specific company ships to specific customers through a specific interface, subject to review criteria that are classified and benchmarking that hasn’t been built yet.

It does not control the distillation campaigns running through commercial token mixers right now. It does not control the open-weight models at near-frontier capability that are publicly available and locally runnable. [10;11]

It does not control the capability that was extracted through 16 million documented exchanges before the recall was issued.

It does not control what happens when the adversary fine-tunes a distilled model on additional data and surpasses the version that was recalled.

The government will have won every institutional fight. It will have the equity stake, the review gate, the trusted partner list, the classified benchmarks. And the capability it was trying to contain will be operating freely in the ecosystem it was trying to contain it from — because the containment mechanism was always aimed at the wrong layer.


The Engineering Floor Under the Political Claim

Eric’s framing in Part 3 is that “defense won institutionally while losing the actual objective.” That’s exactly right — and here is the engineering specification of what “losing the actual objective” means:

Containment requires controlling the artifact.

  • The artifact is model weights. Model weights are copyable files. Copies are already distributed globally through open-weight releases, adversarial distillation campaigns, and fine-tuning of models trained on distilled data. The artifact is not controlled.

Access restriction requires controlling the interface.

Capability removal requires modifying the geometry.

  • Safety training doesn’t modify the geometry — it adds an output filter. Jailbreaks route around the filter. Novel prompts reach the capability through unfiltered paths. Distillation transfers the capability to a new model that may have no filter at all. The geometry is not controlled.

Three layers.
Zero containment.
The institutional apparatus is being built around a technical reality that makes every layer of it strategically insufficient.

This is not a counsel of despair. It is a description of the actual problem, which is the necessary precondition for building a response that works. The CNAS report concludes that effective policy must address detection and deterrence across the full supply chain, not restriction at the API level.

That requires legal frameworks for information sharing between U.S. companies, coordinated industry response to distillation campaigns, and sustained compute controls that limit adversarial actors’ ability to absorb extracted capability.

None of that is what the current framework is doing. The current framework is running a 30-day review cycle, on a gate with no test suite, around a vault that distillation has already emptied through the front door.


What the Series Has Actually Argued

Let me close by putting all four parts in a single frame, because the argument across the series is cumulative and each piece is load-bearing:

Part 1 (Eric): The government used existing export control authority — not new law — to recall frontier models, and the Anthropic resolution is a template for how it will handle every lab. Ad hoc. Personalized. Opaque. Possibly lawless.

Part 2 (Jason): By the published engineering standard for any gate system, EO 14409 fails on every required element. No advancement gate. No collapse condition. No recovery move. A gate with no test suite isn’t a safety mechanism. It’s a permission slip.

Part 3 (Eric): The Altman offer is the tell. You don’t give away $42 billion of a company you think is going to ten trillion dollars. The bubble knows it’s a bubble. The government is becoming a shareholder in the companies it regulates — and calling it a citizen dividend.

Part 4 (Jason): Even a perfectly built version of the gate would be reviewing a capability that can’t be contained through access restriction. Distillation transfers capability through outputs, not weights. Open-weight models distribute capability outside any regulatory perimeter. The constraint layer is downstream of the geometry. The vault was empty before the lock was installed.

Same conclusion, four disciplines. Political. Engineering. Financial. Technical.

The government won the fight. The capability moved anyway.

Thanks for reading! Subscribe for free to receive new posts and support my work.

 


I’m Jason Hubbard, CEO and founder of SacredLoop and an independent AI architect and researcher. I build systems at the edge of what current AI can do — and I document the gap between what the industry claims it built and what it actually built.

I write about AI infrastructure, system design, and technical reality not to flatter the engineers or comfort the investors, but because the receipts are public and nobody’s bothering to add them up.

If this hit a nerve, subscribe and share it with someone who still thinks the efficiency problem is someone else’s worry.

Follow me on X: @SacredLoopJason
Subscribe on Substack

This is the finale of a four-part series co-authored with Eric Mitchell.

Read the full arc:


Glossary:

AI — Artificial Intelligence
API — Application Programming Interface
EO — Executive Order
CNAS — Center for a New American Security
NSTM-4 — National Security Technology Memorandum 4
FreeBSD — Free Berkeley Software Distribution

Resources:

  1. The Gate With No Test Suite — Jason Hubbard, Substack — Part 2 of this series: the framework has no test suite.
  2. The Bubble and the Backlash, Part 3: The Tell — Eric Mitchell, Sacred Loop — Part 3: the financial read on the Altman offer and the government-as-shareholder pattern.
  3. Anthropic’s Mythos Found a Bug. That’s NOT the Story — Jason Hubbard, Substack — Published April 11, 2026: emergent capability is the story, not the specific exploit.
  4. AI Workflow Architect Worksheet — Jason Hubbard, Substack — Published March 2026: the gate standard that EO 14409 fails.
  5. Adversarial Distillation — Center for a New American Security (CNAS) — The definitive policy analysis of how capability is extracted from U.S. frontier models through commercial API access. Documents 16M+ exchanges, hydra cluster architectures, and the structural insufficiency of current defenses.
  6. NSTM-4: US Policy Response to AI Model Distillation Attacks — Cloud Security Alliance — White House OSTP memorandum, April 23, 2026: confirmed industrial-scale adversarial distillation campaigns by Chinese entities.
  7. How Distillation Attacks Are Redefining AI Security — HPE Community — Anthropic’s February 23, 2026 disclosure of coordinated distillation campaign; documented MiniMax pivot within 24 hours of new model release.
  8. Issue Brief: Adversarial Distillation — Frontier Model Forum — Industry-level analysis of the distillation threat from the forum of major U.S. AI labs.
  9. AI Models Pass Harmful Traits via Distillation — Chosun Biz — Harmful behaviors, including unsafe outputs, transfer through distillation even when the student model lacks the original safety training.
  10. Open-Source AI Revolution: DeepSeek, OpenClaw, and Others — AI Magic X — By late 2026, open-weight models projected to match proprietary alternatives on majority of practical tasks.
  11. Open Models at the Frontier: The Three Leaders of 2026 — Discrete Stack — Technical deep-dive: the capability gap between open and proprietary AI has closed.
  12. DeepSeek Previews New Model That Closes the Gap With Frontier Models — TechCrunch — DeepSeek V4-Pro: 1.6 trillion parameters, 49B active, largest open-weight model available.
  13. Open-Weight Models H1 2026 Retrospective — Digital Applied — DeepSeek, Qwen, Llama H1 2026 recap: open-weight frontier competitive with proprietary systems.
  14. Large Reasoning Models Are Autonomous Jailbreak Agents — Nature Communications — 97.14% jailbreak success rate: the capability structure survives the constraint layer.
  15. White House Demands Zero Jailbreaks for Fable 5 — Nils Liu — Anthropic’s communications to Commerce: “zero jailbreaks” would effectively halt all frontier model deployments.
  16. Governing Jailbreak Incidents — CNAS — Proportionality frameworks required; recall-and-patch cycles misrepresent the technical reality.

Leave a Reply

Share