When a feature fires, what do you actually know?
Reading the white-box interpretability claims in the Claude Mythos Preview system card In this article , I explain what we actually know when a feature fires. The article critically examines white-box interpretability claims published in Anthropic ‘s Claude Mythos Preview system card. I look at a specific claim the card makes about “concealment features fired […]