Developers integrating open source AI models into their applications are increasingly falling victim to "open washing," assuming the tools they adopt offer full transparency and modification rights. As the AI industry rapidly redefines traditional software terminology, the distinction between genuinely open source systems and mere open-weight models has become a critical legal and operational boundary. The confusion threatens to dilute software freedoms that have protected developers for decades.
Speaking at Open Source Summit Europe in Prague, Peter Farkas, CEO of Percona and co-creator of FerretDB, issued a stark warning to the developer community. He urged the industry to stop using the terms "open weight" and "open source" interchangeably. "'Open source' is only going to remain 'open source' as long as we preserve the actual meaning and the freedoms that are behind open source, and open weights are not providing that," Farkas explained.
The Dominance of Open-Weight Models in Production
The distinction is not merely academic; it directly impacts how enterprise applications are built and scaled today. Open-weight models have become a foundational component of production AI. In August, these models accounted for 56% of tokens processed through Vercel's AI Gateway. Furthermore, they represented 60% of US-originating token consumption on OpenRouter, with Chinese-developed models driving the majority of that volume.
An AI model's weights are the numerical parameters produced during training, which encode the patterns the system has learned. Releasing these weights allows developers to download and run a model on their own infrastructure. However, this access does not include the original source code or the underlying training data. Farkas noted that because developers only receive the output of the training process, labeling these models as open source AI is fundamentally inaccurate.
Open weights answer 'can I run this?' Open source answers 'can I trust this, improve it, and build the next thing on top of it?'
- James Landay, Stanford HAI
The spectrum of transparency varies wildly across the industry. Farkas pointed to China's DeepSeek as a prime example of the current confusion, noting that its models are widely described as open source despite the company only releasing model weights. Conversely, Xiaomi recently launched its MiMo-V2.6 models with a much higher degree of transparency. Xiaomi released more than 7,000 reinforcement-learning task environments, alongside its RL training code and technical documentation, offering deep visibility into the fine-tuning process.
The Open Source Initiative Reopens the Debate
While Farkas stressed that open weights are highly valuable for experimentation and production, he warned that accepting a looser definition of open source for AI could weaken the standard for all software. He cautioned that companies might eventually label software under an Apache 2.0 license even if it carries hidden regional restrictions, fundamentally breaking the open source social contract.
The Open Source Initiative (OSI) attempted to standardize this terminology by publishing its first Open Source AI Definition in 2024. However, the criteria - particularly regarding how much training data must be disclosed - remain highly contested. Duane O'Brien, the OSI's executive director, acknowledged the ongoing friction during the summit. "The criticisms that have been lodged against open source AI definition 'one' are valid," O'Brien stated, confirming that the organization is reopening the conversation.
To build consensus, the OSI has launched an Open Source AI Fellowship, appointing Gabriel Toscano as its first fellow. The two-year program will explore potential revisions to the definition and host a series of community "open source salons" to ensure the final standard reflects the actual needs of developers and researchers.
The Compliance Nightmare Hiding in Plain Sight
The OSI's willingness to revise its 2024 definition is a necessary retreat, but the current ambiguity is a ticking time bomb for enterprise compliance. The Vercel and OpenRouter statistics prove that developers are aggressively deploying open-weight models into production environments. If the industry continues to casually label these as open source AI, procurement and legal teams will inevitably clear them for enterprise use under false pretenses.
When a company adopts traditional open source software, they know exactly what they are inheriting. With open-weight models like DeepSeek, the lack of training data transparency means organizations are blindly absorbing potential copyright infringements and biased data structures. If the OSI fails to establish a rigid boundary that explicitly excludes mere open weights from the open source definition, the resulting legal fallout will not just harm AI startups - it will permanently erode the trust that enterprise companies place in the entire open source ecosystem.