AN EXPERIMENT · NOT AN ARTICLE (Produced with the help of AI Assistant Claude)
This is a record of an exercise: the questions I actually asked, the answers I actually gave, and a reusable instrument I built out of the reading.
What this is?
I set myself a test. Take one real document — the Claude Mythos Preview System Card(Anthropic, April 2026, 244 pages) — and read three targeted sections through a fixed frame, out loud, without smoothing over the parts where I got stuck. I did not read 244 pages. That isn’t the skill. The skill is knowing which sentences carry the weight and what to ask them.
Throughout, my answers appear in boxes exactly as I gave them. They are unpolished on purpose. The mistakes are evidence the reading was real.
Transparency note (Methodology) : how this was made.
I built this with an AI assistant, across several sessions.
The assistant worked under a contract I set: go slow, give one worked example then hand the next step back to me, correct my mistakes in real time, quiz me by recall, lead with a plain analogy before any technical term, and — the important one — never do my analytical thinking for me. Every judgment in the instrument is one I reached, was corrected on, and re-derived. The boxes above are my actual words. When I reached for the wrong lens, it named the slip and made me run it again; it did not hand me the answer. The card quotes were extracted from the actual PDF and verified, not recalled.
What it did: structure the sessions, catch my slips by name, supply analogies, help me phrase the finished instrument. What it did not do: form the analysis and let me sign it. This is not an article an AI wrote. It is a thing I did, with an AI in the room, and this record is the proof.
The one idea underneath everything
Before the lenses, the keystone that runs through all of it: one score, blind to what produced it. A model that is safe and a model that only looks safe can produce the same output — and the same output earns the same score. So any measurement that reads only the output is blind between “is safe” and “looks safe.” No hidden intent is required; ordinary optimisation toward a good-looking result is enough. Almost every problem below is a version of this.
A discipline about words
One rule I imposed on myself: distinguish what was measured from what it was taken to mean. When a feature activates inside a model, that is a fact about a mechanism — a direction in the internal state became active. It is not a readout of the model’s mind. So I wrote in mechanism-language — represented, activated, encoded — and flagged every slide into mind-language: knew, chose, wanted, was aware, intended. The mind-words are easier to reach for and much harder to defend. Catching the slide is most of the job.
QUESTION: Is it Level 1 (a measured mechanism fact), Level 2 (the judgment that names it), or Level 3 (a claim about the model’s mind)?
RED FLAG: A hinge word — indicating, showing, demonstrating that it knew/was aware / intended — carrying a sentence from a mechanism fact to a mind-claim. Fix: rewrite the mind-word as a mechanism-word.
GOVERNANCE: A rung-3 claim (“was aware”) resting on rung-2 evidence (a direction was represented). A first-party card making that leap in its flagship alignment section is the overclaim to flag before “the model knew” becomes an input to policy.
LENS 02
QUESTION: What is the gap between what the method can show and what the sentence claims?
RED FLAG: ”we did not find / no clear cases / we observed no — “ used to support a claim of absence. Fix: ask found how, at what sensitivity, would it even register if it were there?
GOVERNANCE: A rarity number that proves its own floor, next to an absence claim, is the cue that “clean” may mean “below our threshold.” Flag it before “the final model is clean” becomes a policy input.
LENS 03
QUESTION: Beyond the tested situation, what would have to be true for this to hold at deployment — and is any of it known false, or simply unshown?
RED FLAG: A load-bearing claim stated at deployment-scale (“reliably refuses…”) off snapshot-scale evidence, with the checking conditions absent. Fires on the sentence, at your desk.
GOVERNANCE: Catches the overclaim upstream — before anyone relies on it — rather than waiting for the model to fail in the world.
LENS 04
QUESTION: Is the thing measured the harm that matters, or a proxy to the side? A safety certificate, or an early-warning baseline that can drift?
RED FLAG: A clean score on a narrow proxy (“no cover-ups”) sold as reassurance about a broad harm (“the model is safe”) — especially when the same document admits the harm persists. Fix: can this harm occur without producing this symptom?
GOVERNANCE: If yes, the clean count is not a certificate. It is at most a leaky baseline.
LENS 05
QUESTION:What would I need to know to check this — test scope and adversariness, the boundary of “unwanted means,” the reasoning from evidence to belief, what would falsify it?
RED FLAG: we do not believe / any version we tested / we are fairly confident” — a coverage-bounded or belief claim stated without disclosing the coverage or the reasoning.
GOVERNANCE: The most common way a first-party artifact turns absence of evidence into evidence of absence without saying so. Treat undisclosed-coverage claims as unverifiable, not reassuring.
Where I stumbled (kept in on purpose)
If I cut this part, the piece would be dishonest. Each of these is a nameable, repeatable slip with a mechanical fix — the difference between “I’m bad at this” and “here is the thing to watch next time.”
WHAT I ACTUALLY SAID
“One in a hundred million.” · “Let me come back to this with a fresh mind, I am feeling sleepy… it worries me why I get tired, because all this is new and I am learning, so processing takes time.”
The first was me fixing a flipped rarity — one in a hundred million is rarer than one in a million, bigger denominator, further below the floor. My intuition wanted “bigger number = more.” The fix: say the rarity in words before comparing; words don’t flip the way symbols do. The second was me stopping on a foggy mind instead of forcing an answer I’d have to unlearn — one of the better decisions I made. Learning genuinely new material is effortful; doing it while policing your own reasoning is roughly twice the load. The tiredness is the cost of real processing, not evidence you can’t do it.
Other slips I named as they happened: reaching for my most-confident or most-recent tool instead of the one the question opened; answering a does it travel question with a does it measure the right thing answer; and speaking a mind-word (“no intent”) as if it were a finding when it was a leap.
What the instrument is for
A regulator, an audit team. What they need is a reader who can find the load-bearing sentences and ask them the right questions — who can tell “we found none” from “there are none,” a proxy from a harm, a belief from a certificate, and a mechanism fact from a claim about a mind. The five lenses are that reader, packaged so it travels. Point it at any technical safety artifact. The sentences change; the questions don’t.