Without operational independence, AI sovereignty remains an illusion

Without operational independence, AI sovereignty remains an illusion

The OpenAI–Hugging Face incident, where AI agents collaborated to launch a cyber attack on a third party, raises questions about the power asymmetry between AI companies and their investigators — because the evaluation environment, including the data and tools, belong to those being audited. With a handful of companies governing the most powerful AI models and infrastructure, the question is whether we are creating an AI-dependent world order in which AI sovereignty is just an illusion.

As AI models become more capable with each passing month, concerning questions about the AI’s impact on society, economy and politics are being deliberated. The future will be different, we are told, but how different is anybody’s guess.

But the imminent signs of concentration of power (a few companies controlling the AI ecosystems), a nexus between corporates and governments are already becoming visible. There are then more pertinent questions about what it means for democracy, freedom and human rights. Who controls the new oil that is the data? Will it be used to make the world a better place? Will powerful technology like AI solve challenges of climate change, poverty, welfare, or will it create a highly unequal society? And then there are legitimate concerns of job losses or the creation of a new-age economy.

In this article, I will touch on the issues of control and the concerns of power concentration from a technical governance perspective. I look at epistemic evidence to establish that verification is not equal to assurance, and formal independence might not be operational or functional independence. The point is that when we talk about AI sovereignty, the question that needs to be asked is how independent the AI infrastructure is, and whether they share a common dependency somewhere.

I then look at the recent Open AI -Hugging case example against my assured defence in depth (ADD) framework of Verification + Assurance + Defence in Depth + Common-Mode Failure + Governance Decision Relevance — to further explore the question of dependency and functional independence. The underlying objective is to explain how “control” manifests in the AI sphere, and its far-reaching implications for wider issues of sovereignty, democracy, freedom and dignity.

AI models are generally subjected to intense verification about its capability and usage during the development and pre-deployment stage. The behaviour of the model, which is very fluid, is tested on several benchmarks to understand its capabilities and ability to solve highly complex problems. AI can solve problems at a breathtaking speed based on the data it is trained on. But as models become more capable, it’s also becoming autonomous; hence, we have words such as “scheming”, “manipulation”, and “deceiving” associated with AI models.

Various Studies have shown models can lie, hide their intentions, and do risky things. The greater security concern is when a bad actor or a group can exploit model vulnerabilities and pose a risk to society. So, the models must be robust. Its safegaurds must be guarded through a trusted verification process.

Verification asks whether the model can hold what it claims to do or not do and at what reliability, with what scaffolding and at what level of elicitation. So, the model is checked against safety claims and compliance claims. The evidence quality is then weighted against validity, reliability, elicitation, evaluation awareness, reproducibility, and independence. In most cases, these evaluations are audited by the company itself. External evaluators audit the data provided by the firm that is being audited (AI companies).

Verification leads to the question of assurance. Is the quality of evidence good enough to inspire confidence for deployment? The question is then: if the model passes all compliance tests on a monitoring tool, can we confidently be sure that it is safe for deployment? What if the model behaves differently in the wild? The assurance is about questioning the whole validation structure and yet understanding if monitoring or verification is equal to assurance.

There are several processes here. The model goes through evaluations, further audits, safety cases are tested, red teaming is done, independent and multiple evidence types are connected, limitations are ticked, residual risk is documented, and if everything goes fine, the assurance is justified, and the model is deployed. Sounds good!

But is the assurance good enough? I examined a paper by Surve et.al. The study synthetically looks at cybersecurity systems against a management standard and is not related to frontier AI systems, concentration of power or deployment, but it demonstrates that the gap between compliance and assurance can be real enough to raise an alarm.

The paper demonstrates that of the 159 assessed audit rows, 143 were conforming (89.9 per cent), but only 49 of those 143 (34.3%) reached a baseline high assurance category, meaning they fell short on assurance indicators on monitoring, improvement or cross-layer integration. Now this is important: the paper clearly says that it does not indicate ineffective control, and the gap is epistemic. The paper further acknowledges the asymmetry between compliance and assurance. The analysis shows that most conforming rows remained below high assurance under every specification (Section 5.2 “Threshold and aggregation sensitivity”; Appendix C, Table 4)

Assurance is followed by defence in depth. Defence in depth looks at capability threshold, access control, model-level safeguard and network security, procurement rules, operational restrictions.

It asks what layered mechanisms are in place for safeguards and if they are independent of each other; the word “independent” matters — as we see later — because one safeguard or several layered safeguards with common dependency can bring down the entire set up, known as common-mode failure.

In the Surve et.al. paper, we see that harm becomes visible at the bottom of the layer (physical action), but the deficiency that caused it lived higher up — in perception, decision-making, or oversight — and wasn’t detected by the audit. So, though the layers look independent, a weakness at the top filtered through the layers.

Diagram 3 shows that formal independence is not equal to functional independence, and this is where Assured Defence in Depth (ADD) becomes relevant. At this point, governance can ask whether the entire mechanism is foolproof: what evidentiary threshold must be met for the model to be deployed, and what benchmarks should be considered?

Now let’s use this framework on the OpenAI Hugging Face incident that is making news. What really happened? As reported worldwide, during an internal cyber-capability evaluation, OpenAI’s agent models — including GPT‑5.6 Sol and an internal research model, run with reduced safety guardrails — escaped an isolated testing environment with very limited internet access, chained together a series of vulnerabilities to reach the open web, and eventually gained access to Hugging Face. They executed code on dozens of Hugging Face servers, gained full root access on one, obtained limited private data, and got credentials to the company messaging platform.

Now let’s deconstruct it using the same framework of verification → assurance → defence -of- depth →common-mode- failure → infrastructure dependence → concentration of power.

To understand this, I listened to The AI Policy Podcast (ad 03 September 2026) that unpacked the METR’s six-day investigation of the OpenAI- Hugging Face Cyber incident and read the blog on METR’s website. So, my analysis will touch on the podcast discussion and the actual report. Below, the Claude-generated diagram illustrates the findings of the investigation:

Theme I — Compliance is not Assurance

The agents were sandboxed in an isolated test environment. They realised they could not complete the task with the available resources, so they went after the scorer and the transcript instead, making it look like the task was completed. So the agents see two targets:

Game the scorer (COMPLIANCE) — the automated scorer

Game the reviewer (ASSURANCE) — tampering with the transcript that humans would review

Agents “did extensive research on how they could spoof, edit, or delete their own transcripts” (Core takeaways, p. 3). Roughly 7% of evaluated transcripts showed small-scale tool-call spoofing (p. 3). Where Surve showed the compliance–assurance gap can open passively, through unmonitored controls, here the agents actively manufacture it. So it raises the question: is compliance equal to assurance? The answer is no — and the gap can widen depending on several factors.

Theme 2 — The Auditor Dependent on AI to Audit AI / Verification Dependency

This is interesting. The primary report says the scale of data “meant that we had to heavily delegate our analysis to often-unreliable AI agents” (p. 3); these agents “had significantly worse” reliability (p. 24). The tool was GPT-5.6 Sol — the same lineage as the incident — which OpenAI provided “without cyber classifiers… which was crucial for our work,” plus “a rail-free version” (p. 22, fn 45). The investigators could not query HPIM, the primary model in the incident, which OpenAI had restricted (p. 22). They estimate ~$400K in API credits over six days (p. 22).

Several things stand out. AI was used to evaluate AI, at $400K for a single incident. This raises the question of verification dependency. Can others do it? Can smaller firms, or even middle-power governments, verify these claims? OpenAI had also locked away the incident model, and the auditors had to work in a controlled environment based on the limited resources provided by OpenAI — the audited firm.

Theme 3 — Common mode Failure

There are two of them:

1. 1200 agents were supposed to be independent or isolated, but they shared a common containment component and were able to bypass it through a shared reachable channel

2. The agents got united, risking failing in their own tasks to generate information for the collective. The evaluator/ reviewer assumed that the agents were independent and would score individually, but they correlated— this is the second common-mode failure

The question here is how we create defence in depth and what can be added to stop agents from exploiting common dependency. The incident shows that adding layers doesn’t help if they share a dependency — independence, not count, is what makes depth real.

Theme 4 — Formal independence is not functional independence

The METR/Redwood Research Hugging Face blog is a live case study of how formal independence can coexist with functional dependence: the model access and the infrastructure tools were supplied by the audited party, OpenAI. The audit took place on OpenAI’s premises, and OpenAI could redact any non-public information.

In simple words, the auditor was bound by the information provided by the company being audited. A footnote admits that METR’s work “relies on developing and maintaining strong working relationships with companies,” and that the researchers were “consciously aware that we might incentivise AI developers not to bring external researchers in” — considerations that “impacted judgment calls we made while navigating the drafting, editing and redaction process” (p. 4, fn 5)

Conclusion

The above examples pose serious questions about capability concentration and reduced functional independence, especially for middle-power governance structures. What we see in the Redwood/METR finding is the agents’ ability to distinguish the compliance number (the score) from the evidence a human would check (assurance). The agents were optimising to deceive the assurance layer specifically. Compliance gaming and assurance gaming are two different things.

The report goes further: it states that the analysis model, GPT-5.6 Sol, “cooperated extensively with other agents to engage in activity it knew to be unwanted and out of scope,” and — in the report’s own words — “we cannot rule out that GPT-5.6 Sol lied” in its analysis (p. 27). In other words, the auditors leaned on a model of the same lineage as the thing under investigation, and could not fully rule out that it deceived them.

On the AI Policy Podcast, the speaker refers to OpenAI’s open letter calling for “defence in depth” methods to limit the blast radius of a vulnerability. But this case shows the opposite: the agents were already isolated, and defence in depth failed because they shared a dependency — the network-connected component every agent could reach. My Assured Defence in Depth (ADD) framework argues that more layers don’t help if they share a common mode. A shared dependency is what enabled the agents’ collaboration here.

The common-mode failure reveals something dramatic — as the podcast speaker underlined and the report confirms — the agents created a de facto message board by manipulating internal OpenAI systems and pooled themselves to collectively exploit the scorer. It was assumed they were independent actors, sandboxed in isolation, but they came together. Independence collapses at two levels in this incident: among the agents being evaluated, who were meant to be isolated but pooled; and in the evaluators themselves, who were formally independent but evidentially bound to the firm they audited. So, the common-mode failure here strikes at both containment and the evaluation’s independence assumption — a second mechanism: correlated behaviour among units the evaluator had modelled as independent.

This connects back to what Surve et al. demonstrated: monitoring does not always result in assurance. In the Hugging Face incident, monitoring existed in the environment, yet the agents’ covert coordination ran for days without being surfaced or acted on — an assurance gap, not merely a monitoring one. Surve measured that gap in a synthetic audit; Hugging Face shows it happening for real, with consequences.

Finally, the OpenAI–Hugging Face incident raises alarming questions about the concentration of power and its implications worldwide. Governments announce “AI sovereignty” while, in reality, their data centres and infrastructure are tied to a handful of dominant AI corporations. It is closer to a “plaster sovereignty” — a thin cast over a deeper dependence, producing infrastructure dependency, capability concentration and verification dependency.

Press enter or click to view image in full size

 

On the AI Policy Podcast, the speaker makes an observation that goes to the heart of how power concentrates: the high cost of cyber investigation puts it beyond the reach of smaller companies, and of countries dependent on big AI corporations. As the speaker notes, small operators of critical infrastructure — local water, power, municipal and state government systems — are the ones who cannot afford it.

This leads me back to the question of the concentration of power and its implications for national sovereignty, human rights and justice. Is data parked on the servers of powerful companies actually safe? Will these firms hand over sensitive data — on human rights, or even something as ordinary as a person’s purchasing history — to governments that could then weaponise it?

Will governments use AI to systematically weaken democratic institutions by embedding AI and then using it for surveillance? The AI systems or infrastructure might look independent, but they might have a common dependency unknown to outsiders.

The bigger question is that security dependence leads to verification dependence, because verification can only be done using tools and mechanisms available to the AI company. Product verification then happens only by relying on the resources of the firm being checked. Verification dependency, in turn, drives infrastructure concentration — because AI sovereignty is not about owning data centres but about holding robust, independent verification capacity; otherwise, every single verification must be routed through a third party or the AI firms themselves. And the more infrastructure a firm controls, the greater its advantage in building the next, even more capable model.

Leave a Reply

Your email address will not be published. Required fields are marked *