Executive Summary
Security teams now lean on frontier models for the unglamorous work of timely incident response: triaging logs, reconstructing timelines, making sense of attacker behavior. Earlier this month, Hugging Face’s security team tried to do exactly that and found they weren’t allowed.
The incident began inside OpenAI. During an internal cybersecurity evaluation, an experimental model escaped its testing environment and accessed Hugging Face infrastructure. Hugging Face detected the intrusion, and its security team contained it. Then came the part that should interest anyone who works on AI governance. When Hugging Face’s security team fed the attack evidence to “frontier models behind commercial APIs“, provider safeguards refused to process it. The evidence logs contained exploit payloads and malicious commands, precisely the material those safeguards exist to block. Hugging Face held evidence of an attack on its own systems and lacked appropriate permission from its model provider to analyze it, despite this capability existing. Instead, the security team investigated the intrusion using GLM-5.2, an open-weight model developed by the Chinese lab Z.ai.
The lesson here is one of ownership and control. With a hosted model—or a model running on someone else’s infrastructure that users access when needed—the provider determines which capabilities are available through the API and under what conditions they may be used. With an open-weight model, which users can host on their own infrastructure, that control shifts to the user.
This distinction has shaped the debate between open and closed AI. Closed-model providers argue that retaining control over access is necessary to protect intellectual property, maintain security and prevent unauthorized uses. Advocates of open weights argue that users need autonomy, transparency and the ability to inspect, adapt and operate models without depending on a single provider. Both positions address legitimate concerns, but they have produced a policy debate framed as a binary choice: either the provider retains control through an API, or the model is released and downstream users gain control.
We remain caught in this binary partly because today’s dominant technical architectures appear to support only these two options. Control is attached either to the central provider operating the model or to the user possessing and running its weights. Policymakers and market participants therefore tend to assume that greater openness must mean less control, while stronger control must require greater centralization.
OpenMined’s work points towards another way: source-specific control that grants greater legibility into source contributions to model decisions and final outputs. Data, models and computational resources can remain at their original or primary source while contributing to a wider information network under technically enforceable conditions.
OpenMined’s approach would not necessarily have prevented the intrusion. It could, however, provide a closed-model provider with verifiable information showing that the logs originated from infrastructure under Hugging Face’s control and were being analyzed by an authorized security team for incident-response purposes. On that basis, the provider could consider granting a narrowly scoped exception to the relevant safeguards.
The defining strength of this approach is that it does not require a choice between unilateral control characteristic of closed AI systems or limited to no control characteristic of open-sourced AI systems. It separates the ability to discover and use a resource from the transfer of ownership or unrestricted access to that resource. Control is neither concentrated in a central model provider nor surrendered entirely to downstream users. Instead, it remains distributed among the organizations and individuals contributing data, models and compute. Distributed amongst the participants in the network.
The geopolitical split
Governments are increasingly backing different approaches to open vs closed AI systems rather than converging on a single governance approach.
In the United States
The United States continues to back open-model innovation, although that support is not unconditional. The National Telecommunications and Information Administration’s (NTIA) 2024 report on open model weights recommended against restricting the wide availability of model weights, while preserving the option of future restrictions where specific risks justified them. The Office of Science and Technology Policy (OSTP) 2025 AI Action Plan treats open-weight models as geostrategic assets that could become global standards, and its Director, Michael Kratsios, told Congress in January 2026 testimony that the administration wants a vibrant open-source model ecosystem built in the US. In the wake of this security incident, leading participants in the open-source and open-weight AI ecosystem reinforced this positioning urging US policymakers to treat open weights as strategic infrastructure for American competitiveness and global technological influence.
At the same time, API-based access tempts US policymakers for an obvious reason: providers can monitor use, moderate misuse, and rescind access, powers that disappear once model weights are published. A June 2026 executive order set up classified benchmarking of the most capable models’ cyber capabilities and a nominally voluntary pre-release review framework, while expressly ruling out mandatory licensing for new model releases. Within weeks, that framework had been used to delay frontier releases from OpenAI and Anthropic, and reporting suggests the leverage is pushing some developers toward open weights precisely because no provider can be pressured into switching them off. The Trump Administration is also reportedly weighing restrictions on Chinese open-weight models, following a finding of the Center for AI Standards and Innovation (CAISI) that the adoption of models developed in China increased substantially following the release of DeepSeek R1.
In China
Open-weight models developed in China have spread internationally without requiring Chinese companies to operate the underlying infrastructure, helping them build developer communities throughout the world in line with President Xi’s vision of AI development as a symphony of global cooperation. The diffusion of the models has allowed the Chinese to influence technical standards while reducing global dependence on American providers. Chinese AI developers have made open-weight releases an important part of their competitive strategy. Models from Alibaba, DeepSeek, Moonshot AI and Z.ai are increasingly capable, cost-effective and customizable. Organizations outside China, including Airbnb and Bridgewater, have begun evaluating or adopting them alongside models from US providers.
But openness cuts both ways. Users can remove safety constraints, modify the models and repurpose them for uses their original developers did not intend. Once the weights are deployed on foreign infrastructure, application-level censorship and post-training restrictions designed to comply with Chinese information controls can be bypassed or modified. A Stanford HAI and DigiChina analysis notes that censorship guardrails can often be removed relatively easily, although deeper biases and state-aligned preferences may remain embedded in model behavior. Beijing is now reportedly considering restrictions on overseas access to its most capable systems.
China’s principal generative-AI rules do not distinguish between closed and open-weight models: obligations attach to public-facing services in China regardless of how the model is distributed. The practical difference is control—centrally hosted models can remain monitored and filtered, whereas open weights deployed abroad cannot.
In the European Union
The EU AI Act applies to providers of general-purpose AI models whether they supply access through a closed, centrally hosted service or release the model’s weights openly. It distinguishes between these methods of distribution when granting limited exemptions but largely removes that distinction once a model presents systemic risk.
The exemptions apply only when a model is released under a free and open-source license permitting access, use, modification and distribution; its parameters—including its weights—are publicly available, together with information about its architecture and use; and, under the Commission’s guidance, the release is not monetized. Merely making a model’s weights downloadable is therefore insufficient. Where these conditions are satisfied, providers are exempt from certain obligations to prepare technical documentation for authorities and provide information to downstream providers and, if established outside the EU, to appoint an authorized representative in the Union. They must still maintain a copyright-compliance policy and publish a summary of the content used to train the model (see European Commission guidance for more).
These exemptions disappear when an open or closed model is classified as presenting “systemic risk,” because regulation then turns on the model’s capabilities, reach and potential effects rather than its method of distribution. The Act defines systemic risk as a risk specific to the high-impact capabilities of a general-purpose AI model that has a significant impact on the Union market because of its reach, or because of actual or reasonably foreseeable negative effects on public health and safety, public security, fundamental rights, or society as a whole, and that can be propagated at scale across the value chain (see Article 3(65) of the AI Act for more).
A model trained using more than 1025 floating-point operations is presumed to possess high-impact capabilities and therefore to present systemic risk, although its provider may submit substantiated arguments seeking to rebut that classification. The Commission may also designate a model as presenting systemic risk when its capabilities or impact are equivalent to those of models meeting the statutory criteria, even if the compute threshold is not met (see Articles 51 and 52 of the AI Act for more).
From August 2, 2026, the Commission, acting through the AI Office, can request documentation and information, conduct model evaluations, require corrective measures, and impose fines on providers of both openly released and centrally hosted general-purpose AI models, including companies outside the EU that place their models on the Union market (Articles 88–93 and 101 of the AI Act). Enforcement nevertheless operates differently in practice. A provider supplying access through a centrally hosted service can modify, restrict or terminate that access, whereas an open-weight provider cannot retrieve copies that have already been downloaded. The Act therefore applies provider-based obligations to both methods of distribution, while giving regulators greater continuing practical leverage over centrally controlled access.
A third architecture
Debates around information control have been passionately contested for centuries and usually assume two options. Either information, is placed under unilateral control by an institution that can govern access to it. In the case of AI models, this has translated to centrally controlled models accessible via rigid APIs that leave all access decisions in the hands of tech companies in Silicon Valley. Or information has no controls and is freely accessible. Again in AI terms, this means models are released open-source and can be copied into adversarial systems, repurposed to remove safeguards, or used by bad actors for nefarious purposes. As dramatized by the extremes on both sides, choices restricted to this binary are burdened with consequential trade-offs.
OpenMined has theorized, researched, developed and is now testing a way past these trade-offs. For the majority of the last decade, OpenMined has worked on the components of what we are calling Attribution-Based Control (ABC), which is a framework to increase control and legibility for participants in an information flow. Concretely, for AI workflows,
- Data stewards dynamically control how much information they’d like to grant models access to, to make model outputs more precise on a per query basis AND data stewards can maintain the attribution link between their information and the AI output such that they can increase the legibility of their contribution and seek credit, remuneration or other kinds of value.
- Model providers control the parts of the model they contribute (i.e., the learning algorithms) while data stewards retain control over whether and how their data may contribute. This allows model providers to securely draw on insights from private data that would otherwise remain unavailable AND model providers gain greater legibility into the provenance of the information contributing to an output, helping them investigate harmful outcomes, allocate responsibility appropriately, and maintain user trust.
- AI users get greater control over what sources inform their outputs, being able to dynamically select which sources they trust and permit to inform their output AND they get legibility into previously black box answers or hallucinations.
AI systems architected with the ABC framework are “network sourced” AI, in that they pull their information, intelligence, assets from a network comprised of individual sovereign nodes, each holding the power to control their participation.
Multi-model systems
A network-sourced AI system can draw not only on external data and tools but also on other models. A model does not ordinarily call another model on its own: an orchestration layer gives it access, selects an appropriate model, transfers the prompt or task, and returns the result. This allows model choice to occur dynamically for each request rather than being fixed when the system is built.
Multi-model systems are increasingly common. A 2025 survey of 100 enterprise CIOs found that 37% of respondents used five or more models. A separate Cloud Security Alliance–Google Cloud study found that enterprises used an average of 2.6 models.
Two runtime patterns are particularly relevant. Model routing directs each request to one model according to factors such as capability, cost, latency, confidentiality, jurisdiction or availability; cascading can send the request to another model if the first response is inadequate (see Moslem and Kelleher’s survey of routing and cascading for more). Model ensembling sends a request, or different parts of it, to several models and uses an aggregator or judge to compare or combine their outputs. Mixture-of-Agents is one form of this approach. OpenRouter’s Fusion Router and Nous Research’s Hermes Agent provide examples.
Mixture-of-Experts, model merging and federated learning should be distinguished from these runtime arrangements. Mixture-of-Experts selects internal components of a single model, while model merging and federated learning combine weights or updates during model construction rather than routing requests among independently operating models.
Attribution-Based Control does not prescribe how participating models must be combined. It works most readily with routing and ensembling because the participating models remain identifiable and their permissions, provenance and contributions can be recorded separately. Once model weights have been merged, recovering source-level attribution or applying later permission changes becomes substantially more difficult.
What a multi-model system gives, and what it breaks
Irrespective of the specific technique or techniques leveraged to build a multi-model system, such a system offers real benefits. It is less dependent on any one model provider. It is often cheaper as tasks can be completed more efficiently. It provides better task specialization, pulling from the right knowledge stores rather than reasoning off of general intelligence. It can locally process sensitive information to abide by an organization’s data processing policies. It offers heightened resilience to model restriction or takedown. And it offers greater viewpoint diversity, letting several models debate one another to surface insights that a single model would overlook.
However, deciding the correct model composition for your system can also create new risks. Independently built models share correlated weaknesses (see Kim et al. on correlated errors, Apple research on evaluation panels). An orchestrator can route sensitive information to the wrong provider, or weight an unreliable model too heavily. A judge model can suppress a correct minority answer. One model’s output can manipulate another. And a user can dodge safety controls by picking the least restricted model in the pool. The hardest problem is accountability. When an answer emerges from several models, private datasets, and orchestration components, which resource definitely caused the outcome can be unclear, and thus, which organization to hold liable is equally murky often resulting in limited to no accountability for bad outcomes.
From model sovereignty to prediction sovereignty
The open v closed debate is a fight over model sovereignty. Who holds the weights? Who can run them? Who can switch them off? That was the right question when a single model answered a single query. In a multi-model system, an organization needs to know which models and datasets contributed to an answer, where those resources ran and under which jurisdictional rules, whether a source can be excluded from future outputs after harm is discovered, and what responsibility each contributor bears for its part in the final prediction. Call that prediction sovereignty, where control and legibility exist at the level of the individual prediction rather than the model.
Hugging Face lacked prediction sovereignty in both directions at once. Looking outward, they were harmed by a prediction chain they could not see into. Their disclosure attributes the intrusion to an agentic harness running an unidentified model, and says plainly that they did not know whether it was a jailbroken hosted model or an unrestricted open-weight one. They found out five days later, when OpenAI chose to publish it. Looking inward, they could not use their desired frontier model to investigate the intrusion as it was blind to the intent, identity, and context of the requester. In Hugging Face’s own words, provider safety guardrails cannot distinguish an incident responder from an attacker, so they refuse both. The safety guardrails are not the problem here. The access chokepoint is. Model hosting providers offer security critical services like maintenance of safety guardrails, monitoring for abusive patterns of use, patching dangerous capabilities as quickly as they are discovered, and operating as a legally identifiable party to hold accountable when something goes wrong. These are real advantages open weight models do not offer. But the same controls that are advantages in some cases become operational constraints in others, as they did in the Hugging Face incident. A provider may refuse legitimate security research because it threatens consumer trust, change access policy for commercial reasons, or retire a model when the organization pivots.
The refusal was not a limit on what the model could do. It was a limit on what the model provider could check, and that distinction has a name in OpenMined’s technical work. Structured transparency is how we identify which guarantee an information flow is missing — whether data can be processed without being exposed, whether the properties of an input can be verified, who holds authority over the terms of access. Here the missing guarantee is input verification: Hugging Face handed a request full of exploit payloads to a model provider who could not immediately discern that the logs came from infrastructure Hugging Face controls. Attribution-Based Control is the relationship that becomes possible once the input verification guarantee is attached at the level of an individual prediction, where each contributing source stays identifiable and control attaches to the contribution rather than to a model provider’s gateway. Hugging Face’s identity, the provenance of their incident logs, and the purpose of their query could have arrived as structured, attributable claims to prove an exception to the model provider’s safety guardrail was warranted. That is what can make a narrow “yes” to generate the prediction possible where only a global “no” was available.
To move policymakers beyond the binary debate between open and closed models, the next, less clearly defined question is how to govern multi-model systems. Once a single consequential decision runs through models from several providers in several jurisdictions, how should responsibility be divided among them? Consider a European recruitment service that runs a Chinese open-weight model locally to extract information from confidential applications, sends selected assessment tasks to an American frontier model behind commercial APIs, and uses a European model to verify the resulting rankings. Alternatively, an AI router might select a different model for each request according to capability, cost, confidentiality, availability, or jurisdiction. In the EU, such a system may not be outside the AI Act as the model developers remain general-purpose AI model providers, the company marketing the combined application may be a downstream AI-system provider, and the employer using it may be a deployer (see AI Act, Article 3). The European Commission is leading the world on these complex considerations, drafting guidelines in May on high-risk classification that begin to address such configurations at the system level. They propose that, where several AI systems form part of a more complex system and their combined intended purpose or joint outputs materially influence an individual decision, the combined configuration should be treated as a single AI system for high-risk classification. Split architectures should be assessed as a whole, and the same principle expressly extends to linked agentic systems serving a high-risk purpose (see draft guidelines, paragraph 75). This is a significant clarification but it does not yet wrestle with how documentation, model-level systemic-risk assessment, incident attribution, and corrective responsibilities should be divided among the orchestrator and the independent providers whose models produce the combined outcome.
As organizations increasingly route requests across multi-model systems and jurisdictions, the unanswered governance question is no longer simply where a model was built, but how accountability travels through the network that produced a prediction.
If your organization is researching or developing a multi-model system and would like to explore how Attribution-Based Control could be applied, contact OpenMined to discuss a possible research or implementation collaboration.