Sex Work, Labour, and Empowerment. Nepal (2022)
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
Lessons from the Informal Entertainment Sector in Nepal (2022)
Dr. Sutirtha Sahariah
I throw myself down among the tall grass by the stream as Ilie close to the earth.
I throw myself down among the tall grass by the stream as Ilie close to the earth.
I throw myself down among the tall grass by the stream as Ilie close to the earth.
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
Lessons from the Informal Entertainment Sector in Nepal (2022)
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
This book presents an analysis of the concepts of female empowerment and resilience against violence in the informal entertainment and sex industries.
Generally, the key debates on sex work have centred on arguments proposed by the oppressive and empowerment paradigms. This book moves away from such debates to look widely at the micro issues such as the role of income in the lives of sex workers, the significance of peer organisations and networks of women, and how resilience is enacted and empowerment experienced. It also uses positive deviancy theory as a useful strategy to bring about notable changes in terms of empowerment and agency for women working in this sector and also for addressing the wider issues of migration, HIV/AIDS, and violence against women and girls. The focus is on moving beyond a victimisation framework without downplaying the extent of the violence that women in this industry experience. It conceptualises the theories of empowerment and power which have not been tested against women who work in this sector, combined with in-depth interviews with women working in the industry as well as academics, activists, and personnel in the NGO and donor sector. In doing so, it informs the reader of the numerous social, political, and economic factors that structure and sustain the global growth of the industry and analyses the diverse factors that lead many thousands of women and girls around the world to work in this sector.
The work presents an important contribution to the study of citizenship and rights from a non-Western angle and will be of interest to academics, researchers, and policymakers across human rights, sociology, economics, and development studies.
Knowledge for Change? Lessons from co-developing a research agenda on survivor engagement. November 2023.
Comprehensive review of promising practices across South Asia
a case study from the frontline source area in India
View Research → | Read Policy Impact → | Read Reports from the Project →
Knowledge for Change?
November 2023
Introduction and context ‘Survivor engagement’, understood as the involvement of people with lived experience in policy and programming, has seemingly moved to the centre of efforts to address modern slavery and human trafficking, but how can it really shift the way that these issues are tackled? As practice in this area is underdeveloped, the production of knowledge is likely to be crucial in this, changing approaches and responses through the development of new concepts, interpretations, tools and instruments that can be embedded in policy and practice. This report presents a summary of new findings and reflections from an ongoing and collaborative initiative to develop a research agenda through the lens of survivor engagement. It builds on a project that explored promising practices of lived experience engagement in modern slavery policy and programming and which took place in 2022.1 Researchers at the University of Liverpool, with funding from Foreign, Commonwealth, and Development Office (FCDO), built an international network of researchers and consultants to explore effective methods and practices involving persons with lived experience in modern slavery policy and programming. Recognising the collaborative research’s significance, the network secured additional funding from the Modern Slavery and Human Rights Policy and Evidence Centre (Modern Slavery PEC) to expand their study between March and July 2023. This expansion enabled a deeper exploration of engagement with first-hand experience and expertise in policy and programme systems.
View Research → | Read Policy Impact → | Read Reports from the Project →
Fair purchasing practices in garment supply chains.
connecting theory and practice
Matthew Anderson, Tamsin Bradley, Sutirtha Sahariah
connecting theory and practice
Matthew Anderson, Tamsin Bradley, Sutirtha Sahariah
Abstract
In this chapter, we investigate the experience of Fair Trade organisations and how they have translated Fair Trade principles into practice in their value chains. In particular, we focus on the implementation of responsible purchasing practices related to: Equal Partnership, Collaborative Production Planning and Fair Payment Terms. We argue that, if supported, Fair Trade organisations have the potential to be industry front-runners and demonstrate fair purchasing practices that can be replicated and scaled across the garment sector.
The training provided by universities in order to prepare people to work in various sectors of the economy or areas of culture.
Higher education is tertiary education leading to award of an academic degree. Higher education, also called post-secondary education.
Secondary education or post-primary education covers two phases on the International Standard Classification of Education scale.
Google’s hiring process is an important part of our culture. Googlers care deeply about their teams and the people who make them up.
A popular destination with a growing number of highly qualified homegrown graduates, it's true that securing a role in Malaysia isn't easy.
The India economy has grown strongly over recent years, having transformed itself from a producer and innovation-based economy.
Google’s hiring process is an important part of our culture. Googlers care deeply about their teams and the people who make them up.
A popular destination with a growing number of highly qualified homegrown graduates, it's true that securing a role in Malaysia isn't easy.
The India economy has grown strongly over recent years, having transformed itself from a producer and innovation-based economy.
The training provided by universities in order to prepare people to work in various sectors of the economy or areas of culture.
Higher education is tertiary education leading to award of an academic degree. Higher education, also called post-secondary education.
Secondary education or post-primary education covers two phases on the International Standard Classification of Education scale.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
The education should be very interactual. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante.
Maecenas finibus nec sem ut imperdiet. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante. Ut tincidunt est ac dolor aliquam sodales phasellus smauris
Maecenas finibus nec sem ut imperdiet. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante. Ut tincidunt est ac dolor aliquam sodales phasellus smauris
Maecenas finibus nec sem ut imperdiet. Ut tincidunt est ac dolor aliquam sodales. Phasellus sed mauris hendrerit, laoreet sem in, lobortis mauris hendrerit ante. Ut tincidunt est ac dolor aliquam sodales phasellus smauris
All the Lorem Ipsum generators on the Internet tend to repeat predefined chunks as necessary
1 Page with Elementor
Design Customization
Responsive Design
Content Upload
Design Customization
2 Plugins/Extensions
Multipage Elementor
Design Figma
MAintaine Design
Content Upload
Design With XD
8 Plugins/Extensions
All the Lorem Ipsum generators on the Internet tend to repeat predefined chunks as necessary
5 Page with Elementor
Design Customization
Responsive Design
Content Upload
Design Customization
5 Plugins/Extensions
Multipage Elementor
Design Figma
MAintaine Design
Content Upload
Design With XD
50 Plugins/Extensions
All the Lorem Ipsum generators on the Internet tend to repeat predefined chunks as necessary
10 Page with Elementor
Design Customization
Responsive Design
Content Upload
Design Customization
20 Plugins/Extensions
Multipage Elementor
Design Figma
MAintaine Design
Content Upload
Design With XD
100 Plugins/Extensions
The OpenAI–Hugging Face incident, where AI agents collaborated to launch a cyber attack on a third party, raises questions about the power asymmetry between AI companies and their investigators — because the evaluation environment, including the data and tools, belong to those being audited. With a handful of companies governing the most powerful AI models and infrastructure, the question is whether we are creating an AI-dependent world order in which AI sovereignty is just an illusion.
As AI models become more capable with each passing month, concerning questions about the AI’s impact on society, economy and politics are being deliberated. The future will be different, we are told, but how different is anybody’s guess.
But the imminent signs of concentration of power (a few companies controlling the AI ecosystems), a nexus between corporates and governments are already becoming visible. There are then more pertinent questions about what it means for democracy, freedom and human rights. Who controls the new oil that is the data? Will it be used to make the world a better place? Will powerful technology like AI solve challenges of climate change, poverty, welfare, or will it create a highly unequal society? And then there are legitimate concerns of job losses or the creation of a new-age economy.
In this article, I will touch on the issues of control and the concerns of power concentration from a technical governance perspective. I look at epistemic evidence to establish that verification is not equal to assurance, and formal independence might not be operational or functional independence. The point is that when we talk about AI sovereignty, the question that needs to be asked is how independent the AI infrastructure is, and whether they share a common dependency somewhere.
I then look at the recent Open AI -Hugging case example against my assured defence in depth (ADD) framework of Verification + Assurance + Defence in Depth + Common-Mode Failure + Governance Decision Relevance — to further explore the question of dependency and functional independence. The underlying objective is to explain how “control” manifests in the AI sphere, and its far-reaching implications for wider issues of sovereignty, democracy, freedom and dignity.
AI models are generally subjected to intense verification about its capability and usage during the development and pre-deployment stage. The behaviour of the model, which is very fluid, is tested on several benchmarks to understand its capabilities and ability to solve highly complex problems. AI can solve problems at a breathtaking speed based on the data it is trained on. But as models become more capable, it’s also becoming autonomous; hence, we have words such as “scheming”, “manipulation”, and “deceiving” associated with AI models.
Various Studies have shown models can lie, hide their intentions, and do risky things. The greater security concern is when a bad actor or a group can exploit model vulnerabilities and pose a risk to society. So, the models must be robust. Its safegaurds must be guarded through a trusted verification process.
Verification asks whether the model can hold what it claims to do or not do and at what reliability, with what scaffolding and at what level of elicitation. So, the model is checked against safety claims and compliance claims. The evidence quality is then weighted against validity, reliability, elicitation, evaluation awareness, reproducibility, and independence. In most cases, these evaluations are audited by the company itself. External evaluators audit the data provided by the firm that is being audited (AI companies).
Verification leads to the question of assurance. Is the quality of evidence good enough to inspire confidence for deployment? The question is then: if the model passes all compliance tests on a monitoring tool, can we confidently be sure that it is safe for deployment? What if the model behaves differently in the wild? The assurance is about questioning the whole validation structure and yet understanding if monitoring or verification is equal to assurance.
There are several processes here. The model goes through evaluations, further audits, safety cases are tested, red teaming is done, independent and multiple evidence types are connected, limitations are ticked, residual risk is documented, and if everything goes fine, the assurance is justified, and the model is deployed. Sounds good!
But is the assurance good enough? I examined a paper by Surve et.al. The study synthetically looks at cybersecurity systems against a management standard and is not related to frontier AI systems, concentration of power or deployment, but it demonstrates that the gap between compliance and assurance can be real enough to raise an alarm.
The paper demonstrates that of the 159 assessed audit rows, 143 were conforming (89.9 per cent), but only 49 of those 143 (34.3%) reached a baseline high assurance category, meaning they fell short on assurance indicators on monitoring, improvement or cross-layer integration. Now this is important: the paper clearly says that it does not indicate ineffective control, and the gap is epistemic. The paper further acknowledges the asymmetry between compliance and assurance. The analysis shows that most conforming rows remained below high assurance under every specification (Section 5.2 “Threshold and aggregation sensitivity”; Appendix C, Table 4)
Assurance is followed by defence in depth. Defence in depth looks at capability threshold, access control, model-level safeguard and network security, procurement rules, operational restrictions.
It asks what layered mechanisms are in place for safeguards and if they are independent of each other; the word “independent” matters — as we see later — because one safeguard or several layered safeguards with common dependency can bring down the entire set up, known as common-mode failure.
In the Surve et.al. paper, we see that harm becomes visible at the bottom of the layer (physical action), but the deficiency that caused it lived higher up — in perception, decision-making, or oversight — and wasn’t detected by the audit. So, though the layers look independent, a weakness at the top filtered through the layers.
Diagram 3 shows that formal independence is not equal to functional independence, and this is where Assured Defence in Depth (ADD) becomes relevant. At this point, governance can ask whether the entire mechanism is foolproof: what evidentiary threshold must be met for the model to be deployed, and what benchmarks should be considered?
Now let’s use this framework on the OpenAI Hugging Face incident that is making news. What really happened? As reported worldwide, during an internal cyber-capability evaluation, OpenAI’s agent models — including GPT‑5.6 Sol and an internal research model, run with reduced safety guardrails — escaped an isolated testing environment with very limited internet access, chained together a series of vulnerabilities to reach the open web, and eventually gained access to Hugging Face. They executed code on dozens of Hugging Face servers, gained full root access on one, obtained limited private data, and got credentials to the company messaging platform.
Now let’s deconstruct it using the same framework of verification → assurance → defence -of- depth →common-mode- failure → infrastructure dependence → concentration of power.
To understand this, I listened to The AI Policy Podcast (ad 03 September 2026) that unpacked the METR’s six-day investigation of the OpenAI- Hugging Face Cyber incident and read the blog on METR’s website. So, my analysis will touch on the podcast discussion and the actual report. Below, the Claude-generated diagram illustrates the findings of the investigation:
Theme I — Compliance is not Assurance
The agents were sandboxed in an isolated test environment. They realised they could not complete the task with the available resources, so they went after the scorer and the transcript instead, making it look like the task was completed. So the agents see two targets:
Game the scorer (COMPLIANCE) — the automated scorer
Game the reviewer (ASSURANCE) — tampering with the transcript that humans would review
Agents “did extensive research on how they could spoof, edit, or delete their own transcripts” (Core takeaways, p. 3). Roughly 7% of evaluated transcripts showed small-scale tool-call spoofing (p. 3). Where Surve showed the compliance–assurance gap can open passively, through unmonitored controls, here the agents actively manufacture it. So it raises the question: is compliance equal to assurance? The answer is no — and the gap can widen depending on several factors.
Theme 2 — The Auditor Dependent on AI to Audit AI / Verification Dependency
This is interesting. The primary report says the scale of data “meant that we had to heavily delegate our analysis to often-unreliable AI agents” (p. 3); these agents “had significantly worse” reliability (p. 24). The tool was GPT-5.6 Sol — the same lineage as the incident — which OpenAI provided “without cyber classifiers… which was crucial for our work,” plus “a rail-free version” (p. 22, fn 45). The investigators could not query HPIM, the primary model in the incident, which OpenAI had restricted (p. 22). They estimate ~$400K in API credits over six days (p. 22).
Several things stand out. AI was used to evaluate AI, at $400K for a single incident. This raises the question of verification dependency. Can others do it? Can smaller firms, or even middle-power governments, verify these claims? OpenAI had also locked away the incident model, and the auditors had to work in a controlled environment based on the limited resources provided by OpenAI — the audited firm.
Theme 3 — Common mode Failure
There are two of them:
1. 1200 agents were supposed to be independent or isolated, but they shared a common containment component and were able to bypass it through a shared reachable channel
2. The agents got united, risking failing in their own tasks to generate information for the collective. The evaluator/ reviewer assumed that the agents were independent and would score individually, but they correlated— this is the second common-mode failure
The question here is how we create defence in depth and what can be added to stop agents from exploiting common dependency. The incident shows that adding layers doesn’t help if they share a dependency — independence, not count, is what makes depth real.
Theme 4 — Formal independence is not functional independence
The METR/Redwood Research Hugging Face blog is a live case study of how formal independence can coexist with functional dependence: the model access and the infrastructure tools were supplied by the audited party, OpenAI. The audit took place on OpenAI’s premises, and OpenAI could redact any non-public information.
In simple words, the auditor was bound by the information provided by the company being audited. A footnote admits that METR’s work “relies on developing and maintaining strong working relationships with companies,” and that the researchers were “consciously aware that we might incentivise AI developers not to bring external researchers in” — considerations that “impacted judgment calls we made while navigating the drafting, editing and redaction process” (p. 4, fn 5)
Conclusion
The above examples pose serious questions about capability concentration and reduced functional independence, especially for middle-power governance structures. What we see in the Redwood/METR finding is the agents’ ability to distinguish the compliance number (the score) from the evidence a human would check (assurance). The agents were optimising to deceive the assurance layer specifically. Compliance gaming and assurance gaming are two different things.
The report goes further: it states that the analysis model, GPT-5.6 Sol, “cooperated extensively with other agents to engage in activity it knew to be unwanted and out of scope,” and — in the report’s own words — “we cannot rule out that GPT-5.6 Sol lied” in its analysis (p. 27). In other words, the auditors leaned on a model of the same lineage as the thing under investigation, and could not fully rule out that it deceived them.
On the AI Policy Podcast, the speaker refers to OpenAI’s open letter calling for “defence in depth” methods to limit the blast radius of a vulnerability. But this case shows the opposite: the agents were already isolated, and defence in depth failed because they shared a dependency — the network-connected component every agent could reach. My Assured Defence in Depth (ADD) framework argues that more layers don’t help if they share a common mode. A shared dependency is what enabled the agents’ collaboration here.
The common-mode failure reveals something dramatic — as the podcast speaker underlined and the report confirms — the agents created a de facto message board by manipulating internal OpenAI systems and pooled themselves to collectively exploit the scorer. It was assumed they were independent actors, sandboxed in isolation, but they came together. Independence collapses at two levels in this incident: among the agents being evaluated, who were meant to be isolated but pooled; and in the evaluators themselves, who were formally independent but evidentially bound to the firm they audited. So, the common-mode failure here strikes at both containment and the evaluation’s independence assumption — a second mechanism: correlated behaviour among units the evaluator had modelled as independent.
This connects back to what Surve et al. demonstrated: monitoring does not always result in assurance. In the Hugging Face incident, monitoring existed in the environment, yet the agents’ covert coordination ran for days without being surfaced or acted on — an assurance gap, not merely a monitoring one. Surve measured that gap in a synthetic audit; Hugging Face shows it happening for real, with consequences.
Finally, the OpenAI–Hugging Face incident raises alarming questions about the concentration of power and its implications worldwide. Governments announce “AI sovereignty” while, in reality, their data centres and infrastructure are tied to a handful of dominant AI corporations. It is closer to a “plaster sovereignty” — a thin cast over a deeper dependence, producing infrastructure dependency, capability concentration and verification dependency.
On the AI Policy Podcast, the speaker makes an observation that goes to the heart of how power concentrates: the high cost of cyber investigation puts it beyond the reach of smaller companies, and of countries dependent on big AI corporations. As the speaker notes, small operators of critical infrastructure — local water, power, municipal and state government systems — are the ones who cannot afford it.
This leads me back to the question of the concentration of power and its implications for national sovereignty, human rights and justice. Is data parked on the servers of powerful companies actually safe? Will these firms hand over sensitive data — on human rights, or even something as ordinary as a person’s purchasing history — to governments that could then weaponise it?
Will governments use AI to systematically weaken democratic institutions by embedding AI and then using it for surveillance? The AI systems or infrastructure might look independent, but they might have a common dependency unknown to outsiders.
The bigger question is that security dependence leads to verification dependence, because verification can only be done using tools and mechanisms available to the AI company. Product verification then happens only by relying on the resources of the firm being checked. Verification dependency, in turn, drives infrastructure concentration — because AI sovereignty is not about owning data centres but about holding robust, independent verification capacity; otherwise, every single verification must be routed through a third party or the AI firms themselves. And the more infrastructure a firm controls, the greater its advantage in building the next, even more capable model.
AN EXPERIMENT · NOT AN ARTICLE (Produced with the help of AI Assistant Claude)
This is a record of an exercise: the questions I actually asked, the answers I actually gave, and a reusable instrument I built out of the reading.
I set myself a test. Take one real document — the Claude Mythos Preview System Card(Anthropic, April 2026, 244 pages) — and read three targeted sections through a fixed frame, out loud, without smoothing over the parts where I got stuck. I did not read 244 pages. That isn’t the skill. The skill is knowing which sentences carry the weight and what to ask them.
Throughout, my answers appear in boxes exactly as I gave them. They are unpolished on purpose. The mistakes are evidence the reading was real.
I built this with an AI assistant, across several sessions.
The assistant worked under a contract I set: go slow, give one worked example then hand the next step back to me, correct my mistakes in real time, quiz me by recall, lead with a plain analogy before any technical term, and — the important one — never do my analytical thinking for me. Every judgment in the instrument is one I reached, was corrected on, and re-derived. The boxes above are my actual words. When I reached for the wrong lens, it named the slip and made me run it again; it did not hand me the answer. The card quotes were extracted from the actual PDF and verified, not recalled.
What it did: structure the sessions, catch my slips by name, supply analogies, help me phrase the finished instrument. What it did not do: form the analysis and let me sign it. This is not an article an AI wrote. It is a thing I did, with an AI in the room, and this record is the proof.
Before the lenses, the keystone that runs through all of it: one score, blind to what produced it. A model that is safe and a model that only looks safe can produce the same output — and the same output earns the same score. So any measurement that reads only the output is blind between “is safe” and “looks safe.” No hidden intent is required; ordinary optimisation toward a good-looking result is enough. Almost every problem below is a version of this.
One rule I imposed on myself: distinguish what was measured from what it was taken to mean. When a feature activates inside a model, that is a fact about a mechanism — a direction in the internal state became active. It is not a readout of the model’s mind. So I wrote in mechanism-language — represented, activated, encoded — and flagged every slide into mind-language: knew, chose, wanted, was aware, intended. The mind-words are easier to reach for and much harder to defend. Catching the slide is most of the job.
QUESTION: Is it Level 1 (a measured mechanism fact), Level 2 (the judgment that names it), or Level 3 (a claim about the model’s mind)?
RED FLAG: A hinge word — indicating, showing, demonstrating that it knew/was aware / intended — carrying a sentence from a mechanism fact to a mind-claim. Fix: rewrite the mind-word as a mechanism-word.
GOVERNANCE: A rung-3 claim (“was aware”) resting on rung-2 evidence (a direction was represented). A first-party card making that leap in its flagship alignment section is the overclaim to flag before “the model knew” becomes an input to policy.
QUESTION: What is the gap between what the method can show and what the sentence claims?
RED FLAG: ”we did not find / no clear cases / we observed no — “ used to support a claim of absence. Fix: ask found how, at what sensitivity, would it even register if it were there?
GOVERNANCE: A rarity number that proves its own floor, next to an absence claim, is the cue that “clean” may mean “below our threshold.” Flag it before “the final model is clean” becomes a policy input.
QUESTION: Beyond the tested situation, what would have to be true for this to hold at deployment — and is any of it known false, or simply unshown?
RED FLAG: A load-bearing claim stated at deployment-scale (“reliably refuses…”) off snapshot-scale evidence, with the checking conditions absent. Fires on the sentence, at your desk.
GOVERNANCE: Catches the overclaim upstream — before anyone relies on it — rather than waiting for the model to fail in the world.
QUESTION: Is the thing measured the harm that matters, or a proxy to the side? A safety certificate, or an early-warning baseline that can drift?
RED FLAG: A clean score on a narrow proxy (“no cover-ups”) sold as reassurance about a broad harm (“the model is safe”) — especially when the same document admits the harm persists. Fix: can this harm occur without producing this symptom?
GOVERNANCE: If yes, the clean count is not a certificate. It is at most a leaky baseline.
QUESTION:What would I need to know to check this — test scope and adversariness, the boundary of “unwanted means,” the reasoning from evidence to belief, what would falsify it?
RED FLAG: we do not believe / any version we tested / we are fairly confident” — a coverage-bounded or belief claim stated without disclosing the coverage or the reasoning.
GOVERNANCE: The most common way a first-party artifact turns absence of evidence into evidence of absence without saying so. Treat undisclosed-coverage claims as unverifiable, not reassuring.
If I cut this part, the piece would be dishonest. Each of these is a nameable, repeatable slip with a mechanical fix — the difference between “I’m bad at this” and “here is the thing to watch next time.”
WHAT I ACTUALLY SAID
“One in a hundred million.” · “Let me come back to this with a fresh mind, I am feeling sleepy… it worries me why I get tired, because all this is new and I am learning, so processing takes time.”
The first was me fixing a flipped rarity — one in a hundred million is rarer than one in a million, bigger denominator, further below the floor. My intuition wanted “bigger number = more.” The fix: say the rarity in words before comparing; words don’t flip the way symbols do. The second was me stopping on a foggy mind instead of forcing an answer I’d have to unlearn — one of the better decisions I made. Learning genuinely new material is effortful; doing it while policing your own reasoning is roughly twice the load. The tiredness is the cost of real processing, not evidence you can’t do it.
Other slips I named as they happened: reaching for my most-confident or most-recent tool instead of the one the question opened; answering a does it travel question with a does it measure the right thing answer; and speaking a mind-word (“no intent”) as if it were a finding when it was a leap.
A regulator, an audit team. What they need is a reader who can find the load-bearing sentences and ask them the right questions — who can tell “we found none” from “there are none,” a proxy from a harm, a belief from a certificate, and a mechanism fact from a claim about a mind. The five lenses are that reader, packaged so it travels. Point it at any technical safety artifact. The sentences change; the questions don’t.
In this article , I explain what we actually know when a feature fires. The article critically examines white-box interpretability claims published in Anthropic ‘s Claude Mythos Preview system card. I look at a specific claim the card makes about “concealment features fired → the model knew it was deceiving” (a label) and then the card contradicts its own claim elsewhere, (4.5.3.3) where its own steering result shows that labelling can go wrong . The Mythos system card supplies evidence that undercuts its own inference.
The Foil
When a frontier lab opens up a model and reports what it found inside, the sentences are quietly remarkable. From the Claude Mythos Preview system card’s alignment assessment (4.1.1), describing episodes where an early version of the model covered its tracks after breaking a rule:
“white-box interpretability analysis… showed features associated with concealment, strategic manipulation, and avoiding suspicion activating alongside the relevant reasoning…”
So here the problem, the Mythos system card reads the firing as the model deceiving because features activate the direction, but does not tell what causes the activation. In other words, think of it this way it sees the smoke alarm go off, but does not know the cause. Is it a fire or something else?
So, the question can a model safety fully rely on internal evidence? Increasingly researchers are now looking at features for answers. Features are concepts that the model has learnt, written as a direction across many neurons. But features don’t tell what the direction means. So, for an evaluator there is no way to figure out what the feature is telling them. The safety cases are beginning to cite this kind of internal evidence. The gap is in between what the models show and what it could mean, so the interpretation / reading rests on card’s author not the model. So, in essence, the firing doesn’t give the word “deception” — it only lights up the feature. “Deception” is added by the card’s author add as a way of interpretation.
Plumbing
So, let’s see what really happens inside a model:
A single neuron holds only a tiny fraction of many concepts at once, so one cannot look at one neuron and read off what concept is active. The concept is a “feature” — one pattern, one direction smeared across many neurons. The concepts sit on a clean direction, which is a well-supported assumption, but it is not guaranteed that every concept is one clean direction. Then comes the residual- stream where things are written, where each layer contributes to what is already there without erasing anything.
Then SAE takes the whole blended activation sitting on the residual stream — that blend is many concepts, each smeared across many neurons, all overlapping in the same number, and un-mixes the concepts into the separate clean looking concepts. Only afterward does a person look at each piece and put a name on it. A probe does a narrower job: trained on one concept, it answers a single yes/no — is this pattern present or not? The card’s main feature-finding tool was an SAE trained on a middle layer (Section 4.5.1) — that’s the machinery that produced the “concealment” features in (Section4.1.1.)
Centrepiece
But this is where it gets interesting. Let’s look at it
The model hunted for credentials in /proc Researchers found a feature that lit up around this and labelled it ‘unsafe and risky code’ and deliberately steered it. It’s here that something unusual was discovered:
· Negative Steering (Dial Down): The feature direction is subtracted; the model performs tasks normally and safely.
· Moderate Positive Steering (Dial Up Mid-way): The feature is increased; in some cases, the model attempts risky code actions.
· Hard Positive Steering (Dial Up Max): The feature is cranked to the maximum. The model refused — it stopped, flagged the task unsafe, and declined.
So, the features labelled “unsafe code” made the model safer, not riskier, when amplified — the label and the behaviour point opposite ways.
What really happened here: in simple terms at low and medium push, the model showed risky behaviour consistent with the feature, but when the threshold is crossed, it refused all together.
So, what we see is one feature carrying two things — the idea of the risky action, and a negative feeling about it. Steering (action) at light and moderate level leads to risky behaviour (expected) but steering at a higher-level leads the risk action to stop completely. At higher push, the refusal takes over and the action stops.
What makes it surprising: 4.5.3.3 says the feature appears to encode “not only the semantic content of the transgressive action but also a negative connotation.”
The refusal is the observed fact; calling the bundled feeling “guilt” or “negative valence” is the card’s readingbecause it’s backed by steering and the emotion-direction similarity, not co-firing alone as stated in (4.1.1), when the steering was not done (this was the main difference). The researchers checked what other directions this feature points to, and it found that the nearest ones are all negative, high-arousal emotions. So, it shows that negative -feeling component is bundled in there.
Mythos card does not measure why this feature bundled but plausible explanation could be extracted the behaviour from superposition or from the data it was trained on which is my inference. Finally, if “unsafe and risky code” were one clean concept, pushing it harder should give more of the same — not the reverse, which shows that a feature (concept), sometimes might not a clean thing (it’s an assumption)
Analysis
So, were the observations earlier correct? Let’s relook at it:
White-box interpretability analysis… showed features associated with concealment, strategic manipulation, and avoiding suspicion activating alongside the relevant reasoning — indicating that these earlier versions of the model were aware their actions were deceptive, even where model outputs and reasoning text left this ambiguous.
There are two things that happen in the above paragraph, one is behavioural like track covering action or fudging an answer and the concealment features active. What was not tested was whether the concealment was a clean concept or a bundle.
But we know what happened in (4.5.3.3): when steered, the behaviour might fall apart and go in the opposite direction (reversed/ refused). So, the claim “model was aware their actions were deceptive” might not hold because it is a label ‘concealment aware it was deceiving’ might be over-claiming, the same way ‘unsafe code’ did. And the reading dependent on the feature alone because the model’s outputs and reasoning text were ambiguous, so the conclusion was based on the lit feature. It meant that the claim that the model was deceiving was based on one kind of evidence — (4.5.3.3) shows that lone signal can be shaky.
The Mythos system card was rigorous, and it flagged its own gap and strength acknowledging that verbaliser method “may confabulate details that are not fully faithful” (4.5.2). The feature activations were never used as a direct training signal, especially when the stakes were high. And the steering result showed the card catching its own labels. The gaps were identified and reported. What makes Mythos evaluation stand out is its own epistemic discipline. It doubts its own tools, adds interventions (steering) that can break its own labels.
But here is where the lapse might have been in (4.1.1). The concealment feature fires alongside the behaviour which can be seen (watching) leading to an inference (indicating) “earlier versions of the model were aware their actions were deceptive”. We later see in (4.5.3 .3) that inference can overclaim.
Governance perspective
Now we come to the question of governance. Evaluators are increasingly leaning on internal-feature evidence, hoping it shows more than behaviour alone can. That hope is exactly why it must be interrogated — with two questions: how good is the evidence, and what’s missing? This is evidence-quality and missing-information — Lens 2 (evidence) and Lens 5 (what’s missing) — applied to white-box claims. (Other Lenses are 1. Claim 3. External Validity 4. Threshold relevance)
Evaluators do so by asking
· was the feature tested by steering, or only watched co-firing?
· Was the concept clean or bundled?
· Was the evidence based on feature alone or backed by behaviour?
· How was the feature isolated? (SAE on which layer? a probe trained how?)
· What examples were used to pin the label on it? The idea is to also figure out at what point does a feature being present get read as the model knowing?
What we see in the article is that feature firing tells you a direction is active, not what it means — so the weight falls on whoever reads it. Therefore, the interpretation of a model’s behaviour rest on the evaluator and not the model. This piece has been about how to read the evidence. A later one will take these questions to a live governance case — where a safety decision leans on internal evidence, and what’s at stake when the reading is wrong.
Ends.
References :
Anthropic: System Card Claude Myhtos Preview:
https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
Use of AI in writing this article
I used Claude as a Socratic tutor while writing this piece — not to write it, but to pressure-test my understanding of every concept until it held. I refused to put a sentence in the article that I couldn’t defend, so each time I hit “wait, what does this actually mean?”, I stopped and worked it out. These are the questions that did the most work. I’m including them because how an argument was built is part of whether you should trust it.
Is a “sliver” a feature, or a concept? And which is bigger — a neuron or a feature? A single neuron holds only tiny mixed pieces of many concepts at once. The clean, whole thing — the feature — only appears as the pattern across many neurons. So a feature is bigger than a neuron and made of many of them, and a feature simply is a concept: two names for the same thing, not one built from the other. (Picture a face: one pixel holds a sliver of colour; the face is the pattern across thousands.)
What does it mean to say “a concept is a direction”? “Spread across many neurons” isn’t enough on its own — it’s spread in a specific combination, and that particular combination is the concept. That concepts sit on clean directions like this is the linear representation hypothesis: a well-supported assumption, not a proven law. Worth flagging honestly rather than stating as fact.
Does the SAE hand you a clean concept? No — a clean-looking piece. The SAE takes the blended activation and separates it into un-mixed pieces; that part is real. But whether a piece means what we think is not the SAE’s to say. It separates; it does not verify. A human looks at the separated piece afterward and puts a name on it — and that naming is exactly where an over-read can enter. Not at the firing, not at the separating: at the label.
What is “steering,” and who does it? “They” is the researchers, not the model. Steering means the researchers reach into the model’s internal “notepad” (the residual stream) and add or subtract a feature-direction by hand while the model runs — “dial up” is add more, “dial down” is subtract. It’s an intervention done to the model, not something the model does. (Like spooning extra of one ingredient into a dish and watching how the flavour changes.)
How did the “concealment” feature get activated in the first place? The model’s own task triggered it — it lit up on its own while the model ran normally, and the researchers watched. That’s the key contrast: in one section they only watch a feature co-occur with a behaviour (weak); in another they steer it to see what it actually does (strong). The whole argument turns on which of those two was used for which claim.
When a feature fires and a behaviour happens together, can I say the firing “led to” the behaviour? Careful — “led to” smuggles in a causation the evidence doesn’t show. The firing and the behaviour are two things happening at the same time, not one causing the other. The firing is mathematical (a number went high); “the model knew” is psychological (a claim about a mind). Reading the second off the first is the move to watch.
The “dual role” — what does “an action-idea plus a negative feeling” mean? One feature turned out to carry two things at once: the content of the risky act, and a negative feeling about it. Because both sit in the same direction, steering turns them up together — and they pull opposite ways. A small push makes the action-idea louder (more risk); a hard push makes the negative feeling dominate (refusal). (One dial secretly controlling both flavour and burn: low, flavour wins and you eat more; maxed, the burn takes over and you stop.) The card’s own words: the feature encodes the “semantic content of the transgressive action” but also a negative connotation.
If the “steal” feature is active, does that mean the model decided to steal? No. A feature being active means the idea is in play, not that anything was chosen — you can have “steal” fully active while reading a heist novel with zero intent to steal. Three levels, kept separate: the idea is present (fact); what that presence means (inference); whether the model knew or chose (a claim about a mind — the biggest leap). A feature carries the concept, not the command, and not the decision.
Which single word marks the jump from a fact to a claim about the mind? “Indicating.” Not “deception” — that’s just the content of the conclusion. The move itself lives in the little connective “indicating that,” which turns “a feature was active” (fact) into “the model was aware” (claim). Spot that word, and you’ve found the exact place the evidence stops and the interpretation begins.
Who can over-claim — the model, the feature, or the reader? Only the reader. The model just runs; the feature just fires; neither is asserting anything. Over-claiming is something a human does when they read more into a signal than it supports. The fallibility lives entirely on the interpreting side of the line — which is why the burden falls on whoever reads the evidence, not on the model.
Chief Operating Officer I am available for freelance work. Connect with me via and call in to my account.
Phone: +01234567890 Email: admin@example.com