The AI industry is facing a fascinating and ironic dilemma, one that many internet users and content creators have encountered before. Tech giants like Anthropic, OpenAI, and Google are now grappling with the consequences of their own actions, learning the hard way that once content is online, it can be used and manipulated in ways they may not have anticipated or approved of.
The Distillation Dilemma
At the heart of this issue is a practice called "distillation." In simple terms, distillation involves using the outputs of one AI model to improve or develop another. This technique has sparked concern among AI companies, as they fear that their hard-earned intelligence and research could be replicated by rivals at a fraction of the cost. However, this fear seems hypocritical when we consider the industry's own practices.
Symmetry in Unfair Use
The irony becomes apparent when we realize that AI companies have been doing the same thing to the internet for years. They've been scraping web content, often without permission, and turning it into their own products, justifying it as "fair use." Now, they're facing the same issue they've inflicted on others. Anthropic, for instance, is accused of extracting intelligence from websites, while simultaneously complaining when its own models are targeted.
Bots and the Cost of Doing Business
The situation is further complicated by the use of bots. AI companies are quick to frame the extraction of intelligence from their models as a cybersecurity issue, pointing to bot "attacks." Yet, they've been employing similar tactics against websites, causing operational costs to soar for those sites. It's a double standard that highlights the industry's lack of self-awareness.
Blurred Lines and Distillation Panic
Even within the AI community, there's disagreement over what constitutes acceptable distillation. Some see it as a benign practice, while others, like Anthropic, view it as an attack. This has led to what's being called "distillation panic," with concerns that Anthropic's aggressive stance may hinder legitimate research efforts.
A Game of Cat and Mouse
As Zilan Qian from the Oxford China Policy Lab puts it, "It's always a kind of a cat-and-mouse game." Once information is out there, it's difficult to control how it's used. This is true for all online content, and AI model outputs are no exception. The AI giants are learning that they can't have it both ways—they can't freely extract intelligence from the web and then expect their own models to be off-limits.
A New Reality for AI Giants
In conclusion, the AI industry is facing a reality check. The practices they've employed for years are now being turned against them, and they're struggling to adapt. This situation highlights the need for clearer guidelines and ethical considerations in the development and use of AI models. It's a complex issue, and one that will likely shape the future of the industry and its relationship with the internet.