The AI Model Distillation Paradox

On Sept. 1, the Department of Justice filed its brief for the artificial intelligence (AI) companies’ right to appropriate and reuse your work—which it called a “statement of interest”—in OpenAI’s copyright battle against the New York Times and other publishers.

The federal government’s first direct intervention in the wave of lawsuits over generative AI training data took the position that feeding copyrighted text into an AI model for training purposes is fair use. OpenAI’s training is “exceedingly transformative,” it argued, and requiring companies to license everything they ingest would “render [the] training of AI models impermissible.”

At stake, the department warned, is all America holds dear: its freedom of expression, its lead in technological innovation, even its national security. By handing a “competitive advantage to foreign adversaries who are not so encumbered,” the department contended, the law would render unto China what should be America’s.

Eight days later, however, the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the FBI were singing a different tune. In a joint cybersecurity advisory report, the government accused six Chinese laboratories of running “industrial-scale distillation campaigns” against American frontier models. Distillation is a method that trains a smaller AI model to mimic the behavior of a larger, more powerful one, allowing it to run on cheaper hardware without sacrificing performance.

In its brief, the government alleged that the malefactors—Chinese firms such as DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.ai—extracted billions of tokens, the basic data units that an AI model reads and generates, from variants of such wholesome American products as Claude, GPT, Gemini, and Grok, “likely with Chinese government awareness.”

This was no occasional shortcut: According to the cybersecurity advisory, “aggressive, malicious, and targeted distillation activities” began at least as early as 2024, and formed the “core” of Beijing’s AI strategy.

When American labs ingest the world’s writing without permission, it’s hailed as transformative innovation. When Chinese labs do the exact same thing to U.S. models, it’s a “malicious and targeted” campaign to “to extract proprietary functionality and reasoning capabilities”.

All of which raises three critical questions:

First, what actually separates training a model from distilling one?

Second, given that both practices rely on taking someone else’s work without consent, why does the federal government regard American and Chinese companies so differently, and is there any legal or rational basis for that difference?

And third, if there isn’t one, how should the law treat the mass harvesting of data—whether the harvested material is written by humans or generated by machines?

by Lisa Klaassen/Lawfare – The AI Model Distillation Paradox

Latest articles

Related articles