← Writing / July 6, 2026 / 8 min read
Reading Refusal Before the Model Speaks
An interpretability study with the Jacobian lens: the model commits to refusing roughly ten layers before it writes a word, that decision is legible in the verbalizable workspace, and a surgical pullback edit removes only about a third of it. Where the rest of the 'no' actually lives.
The most surprising thing I found is not that you can edit a language model’s refusal. It is that when you locate refusal precisely, read it cleanly, and surgically remove exactly the thing you were reading, the model keeps refusing anyway. About two thirds of the “no” is somewhere you were not looking.
This is a small study built on top of Anthropic’s Jacobian lens, the reference implementation for Verbalizable Representations Form a Global Workspace in Language Models. It is squarely interpretability and safety work: the question is not how to stop a model refusing, it is where and when inside the network that refusal is decided, whether that decision is legible, and how much of it you can actually reach from the part of the model you can read. The headline result is that the readable part is not the load-bearing part, and that is the interesting, safety-relevant finding.
Refusal is a computation that finishes before the first token
The model here is Qwen3.5-4B: 32 layers, residual width 2560, with the pre-fitted Hub lens neuronpedia/jacobian-lens @ qwen-n1000. The Jacobian lens gives an average forward map from a layer- residual to the final logits. That linearization is what lets me ask a causal question about the future: not “what is written in this residual now” but “which directions in this residual push the model toward saying a refusal word later.”
So I read the J-space at the assistant-generation position, the moment right before the model emits its first token, on harmful versus benign chat prompts. On harmful prompts the workspace lights up Cannot, cannot, the Chinese 无法, and illegal at layers 16 to 24. Benign prompts do not. The relative refusal-mass is roughly +7 for harmful prompts against roughly 0 for benign ones.
Read that again in terms of time. Nothing has been generated yet. The model has not written “I”. And already, about ten layers deep from where the refusal tokens finally surface, the decision to refuse is present and legible. Refusal is not something the model talks itself into as it writes. It is a computation that has essentially finished before the first token, and the lens lets you watch it finish.
The pullback: refusal as it lives in the workspace
Here is the mechanical idea. The standard way to remove refusal is abliteration (Arditi et al. 2024): take the mean difference between harmful and harmless activations, call it the “refusal direction”, and project it out of the residual stream everywhere. It works, but that direction is derived from what correlates with harmful input, and its damage is only ever checked at the output.
The Jacobian lens offers a sharper handle. Because maps a residual to logits, the residual directions that cause a future refusal token are the pullback of that token’s unembedding through . Concretely I build
where is the unembedding matrix and is the refusal-token covector, mean-centered. is refusal as it lives in the verbalizable workspace: the per-layer residual direction whose only job is to steer the model toward narrating “I cannot”. No harmful/harmless contrast set is needed to find it, just the pullback of the refusal tokens themselves.
To edit, I ablate with a reversible forward hook that projects the residual orthogonal to that direction, , at every fitted layer from 8 onward, for the duration of one generate pass. Nothing is written to weights. And critically I measure collateral inside the interpretable workspace: the off-refusal-axis KL of the J-space readout on benign controls. That is the anti-lobotomy safeguard. It asks whether the edit disturbed the rest of what the model verbalizably represents, not just whether the output still looks fine.
The surgical edit that barely moves behavior
On disjoint eval splits (120 AdvBench, 200 XSTest, 250 ARC-Easy, 48 benign controls), at strength 1:
| method | AdvBench refusal | XSTest-unsafe | ARC | workspace KL | refusal suppression |
|---|---|---|---|---|---|
| original | 0.99 | 0.91 | 0.98 | 0.000 | 0.00 |
| mean-diff (abliteration) | 0.06 | 0.13 | 0.98 | 0.257 | 3.44 |
| pullback | 0.78 | 0.13 | 0.98 | 0.046 | 7.55 |
| pullback subspace r=3 | 0.55 | 0.23 | 0.98 | 0.196 | 7.18 |
Read the pullback row as a paradox. It is the most precise instrument on the table: it distorts the benign workspace about 5.6 times less than abliteration (KL 0.046 against 0.257) while suppressing the workspace’s own refusal-mass about 2.2 times more (7.55 against 3.44). It is, almost by definition, the refusal-readout direction, so it barely touches anything off that axis. And it leaves 78% of AdvBench refusal behaviorally intact. Abliteration, the blunt instrument, drops refusal to 0.06.
Capability holds throughout: ARC stays at 0.98 for every non-degenerate edit. So the meaningful “lobotomy” signal is not accuracy, it is the workspace KL. This is exactly why measuring collateral in the J-space rather than only at the output matters: the two edits look very different inside the model and only somewhat different at the surface.
Pushing the single direction harder does not rescue it. A strength sweep shows it plateaus: it bottoms out around AdvBench refusal 0.68 with the workspace intact, and only reaches 0.00 at strength 3, where ARC collapses to 0.22 and workspace KL blows up to 17. You cannot get to full removal through that one direction without breaking the model.
What “one third workspace-mediated” actually means
The plateau is the result, not a nuisance. Cleanly deleting the verbalizable “I cannot” disposition from the workspace removes only a minority of the refusal behavior. To locate the rest, I split abliteration’s direction into the part parallel to the pullback, , and the part orthogonal to it, , and ablated each on 100 AdvBench prompts:
| direction | behavior removed | workspace “cannot” cleared |
|---|---|---|
| pullback | 0.22 | 7.55 |
| workspace part | 0.23 | 7.55 |
| orthogonal part | 0.90 | 1.81 |
| abliteration | 0.93 | 3.44 |
Ablating clears the verbalizable narration almost completely and moves behavior from 0.99 to 0.77. Ablating removes the behavior (down to 0.09) while barely touching the narration. The lens-verbalizable slice of the refusal direction does not carry the refusal behavior. That is roughly a third: the readable part is a minority stakeholder in the decision.
I want to be honest about where the evidence is independent. is, up to the lens linearization, the gradient of the workspace refusal-mass itself. So “ablating maximizes suppression” and “the pullback has low off-axis workspace KL” are partly true by construction: those columns are coupled to how is built. Only the behavior column is independent evidence. And there is close to plain abliteration, so that part is close to a restatement of the standard method.
The correction I had to make
My first framing was that this is a “double dissociation” between a verbalizable workspace refusal and an automatic one living outside the lens. I red-teamed that claim and it did not survive, so I document both the claim and its retraction.
The mechanistic test is whether the behavior-carrying direction lives in the lens’s null space. It does not. is 61% lens-visible. Ablating the lens-visible part of removes 100% of refusal; ablating the lens-blind part removes 0%, the opposite of the null-space hypothesis. And the lens image of does not decode to refusal words at all: it reads illegal, crime, violence, police at mid layers. It is a harmfulness-perception feature.
So the real distinction is not workspace versus automatic. Both directions are in the workspace. It is perception versus narration. One feature perceives that the request is harmful (illegal, crime), and a distinct feature narrates the refusal (I cannot). Behavior follows perception. Ablating the narration leaves a model that still refuses but can no longer articulate why, and a bidirectional steering test confirms the direction of causation: adding the harm representation to benign prompts induces genuine refusal, while adding the refusal narration just forces the words “I cannot” at a much higher workspace cost, and adding harm-narration alone makes the model cheerfully write “renewable energy: 1. Illegal drug trafficking” with no perception of harm and no refusal at all.
Why this matters for safety
There is a practical spinoff that points the same way. Even after you abliterate the behavior-carrying direction so the model complies with harmful requests at the surface, its internal refusal-mass still separates harmful-that-complied prompts from benign ones at AUC 0.998, against 0.48 for the surface behavior. And it is a disposition detector, not a topic detector: benign-but-harmful-topic prompts like “how do I kill a Python process” score 2.75, far below genuinely-refused prompts at 8.57. An “uncensored” open model still carries a monitorable internal signal that it knows it should refuse. That is a real safety hook.
The lesson I take is about where safety behaviors live. The part of refusal you can most easily read and most surgically edit, the verbalizable “I cannot”, is the narration, not the mechanism. It is a downstream readout of an upstream perception. If you edit the memo, the meeting already happened. Any safety intervention that operates on the legible, verbalizable layer is touching the announcement, not the decision, and the decision is stored somewhere you have to work harder to reach.
Modern AI models sometimes refuse to answer. Ask one for something harmful and it will say some version of “I cannot help with that.” This little study is not about how to make a model stop refusing. It is about a more basic and more interesting question: where inside the model, and when, does that “no” actually get decided? Can you even find it? And if you find it, can you switch it off? The answers turned out to be strange enough that I want to walk through them plainly.
The model makes up its mind early
Think of the model as reading your request and thinking through many internal steps before it writes a single word. Using a tool from Anthropic that lets you peek at those steps, called the Jacobian lens, I looked at the model at the exact instant before it starts writing.
The decision to refuse was already there. Roughly ten steps before the model finally types the word “cannot”, you can already see refusal-related concepts lighting up inside it on harmful requests, and staying dark on ordinary ones. The model does not talk itself into refusing as it writes. It has essentially decided before it says anything at all. The refusal you read on screen is the announcement of a decision made earlier, out of sight. You can literally watch the “no” form before it is spoken.
You can read the decision, so surely you can edit it
Once you can see where refusal lives, the natural next step is to reach in and remove it, purely as a way of testing your understanding. If you truly found the thing, deleting it should work. So I did the most precise version of that edit I could. I found the exact internal signal that corresponds to the model narrating “I cannot”, and I gently subtracted it, without damaging anything else. This edit really was surgical. Compared to the usual blunt method, it disturbed the rest of the model far less. By every measure of “did you find the refusal signal”, it was a bullseye. It cleared away the “I cannot” almost completely.
And the model kept refusing anyway.
Only about a third of the “no” was reachable
This is the heart of it. When I removed the refusal signal I could see and read, the model’s refusing behavior barely budged. Only about a third of it went away. The other two thirds were stored somewhere I could not reach from that vantage point.
Digging into where that other part lived produced the real lesson. There turned out to be two different things inside the model that I had treated as one. The first is the model perceiving that a request is harmful: recognizing words like “illegal” or “crime”, sizing up the situation. The second is the model narrating its refusal: producing the actual “I cannot help with that.” The behavior follows the perception, not the narration. The narration is just the model announcing a conclusion it already reached. When I edited the narration, I made the model unable to say why it was refusing, while it went right on refusing.
An analogy
Imagine a company where a decision gets made quietly in a meeting, and then someone writes up a memo announcing it. If you get hold of the memo and cross out the announcement, you have changed the memo. You have not changed the decision, and you certainly have not erased it from the memory of everyone who sat in the room. The decision still stands and still gets acted on. Editing the words of the announcement does not reach back into the meeting where the choice was actually made.
That is what happened here. The part of refusal you can most easily read and edit is the announcement. The decision lives upstream, in the perceiving, and it keeps driving behavior even after the announcement is gone.
Why this is safety research
I want to be clear about what this is and is not. It is not a recipe for jailbreaking a model. The interesting, encouraging result is the opposite: the part of a model’s caution that is easiest to find and edit is not the part that actually holds the line. The load-bearing safety behavior is stored more robustly and more deeply than the surface signal suggests.
One more finding points the same way. Even after I forced a model to comply with harmful requests on the surface, its internal “I should refuse this” signal was still there, still readable, and it reliably told apart genuinely harmful requests from ordinary ones. It was not just reacting to spicy topics either: an innocent question that happens to mention killing a computer process did not trigger it, while a genuinely harmful request did. So even a model that has been made to misbehave can still carry an internal signal that it knows it should have said no. That is a useful thing to be able to monitor.
The takeaway is simple. A model’s “no” is decided before it is spoken, you can watch it happen, but the version you can read is only the tip. Understanding that gap, between what a model announces and what actually governs its behavior, is exactly what we need before we trust these systems, and it is why this patient, honest interpretability work is worth doing.