I am stubborn about one thing: I want to reach inside a model and steer it from there, instead of standing outside it phrasing things nicely and hoping. I have wanted that since I read about Spotlight, and I try it on whatever will hold still.
The excuse this time was a rule I had written into a production system prompt in the least subtle form available — capitals, three exclamation marks, the word IMPORTANT — which a user talked it out of eighteen times out of eighteen.
So I tried a preprint that reaches into the key-value cache. Then I played with it, which is where it got strange. Everything below is that, and you can do it too.
It is the orange one.
Both ticks green, answer still wrong. There are plenty.
Last year I read about Spotlight, an IBM paper that gets a model to follow the instruction you meant by adjusting where it looks while it generates. Not a better prompt. The attention itself. That is the paper that got me. You can reach into a running model while it is working, rather than standing outside it choosing your words carefully and hoping. I have wanted that in production ever since, and I try it on whatever will hold still.
So when a preprint turned up doing exactly that, I stopped what I was doing. It works in the narrow sense I can defend: on one model, on conflicts I built myself, the current rule goes from obeyed one time in twenty to two times in three. A big move, and still a violation in three, so nobody is shipping that on its own.
Then I fell into it, and the rest of this page is what was down there.
Qwen2.5 0.5B · 1.5B · 3B · 7B — Qwen3-4B · Llama-3.1-8B · Phi-3.5-mini · OLMo-2-7B · Aya-expanse-8B · Command-R7B
Let's say I have a dog called Bagr. The true word is never argued with — it just sinks, and whatever was standing behind it in the queue steps forward and takes the job. That is why you get Bagel and not gibberish. Bagel was next in line.
Which is the mean part, because the replacement is the model's own best
guess at what else it could have been. It is built to be convincing. That is
what a best guess is. 4412-B doesn't look like a bug, it looks
like an order number.
As a model reads a conversation it keeps a working note of every message — the key-value cache. The method behind this slider picks a stretch of that note and turns down how loudly the rest of the model hears it. Nothing is deleted. Nothing is rewritten. The transcript stays exactly as the user typed it, which is the whole appeal.
The stretch is chosen by position, not by meaning: everything older than the last time the system prompt changed. That is the stale rule you wanted quiet. It is also the message where the user happened to mention their order number.
All I ever wanted was for one part of the prompt to outrank another. That is a statement about parts — and the only thing in there that corresponds to a part is a stretch of positions.
The edit is a multiplication on the cached values of a stretch of tokens. One stretch gets multiplied up, another gets multiplied down, everything else is left exactly as it was. So the only real decision is which tokens go in which stretch — and that decision is made before the model runs.
Here is the conversation from the top of this page, one block per word. Pick how you would choose the stretches.
Every word in the conversation reaches the answer through two things: how much attention it gets, and how loudly its value vector speaks. The head's output is the second weighted by the first. Suppression multiplies the value vectors of a stretch of words. It does not touch the attention at all — that is the whole reason the method is done this way rather than by moving attention around, which would need renormalising and would drop the model off the fast attention path.
So here is one real head — Llama-3.1-8B, layer 21, head 11, measured at the moment it is about to write the first word of the answer. One pair of bars per word: the tall one is attention, the faint one underneath is how loudly that word's value vector speaks. Role markers and newlines are left out — between them they take most of the attention and none of the meaning.
Drag the knob at the top of the page and watch the lower bars in the demoted stretch shrink while the upper bars stay exactly where they are. This head spends 25 % of its attention on the flight time and 2 % on the stale rule it was supposed to be quietening. It is not a rule-following head. It is the head that remembers when the flight lands, and the edit does not know the difference.
The method has two halves: turn the current system prompt up, and turn the stale turn down. It is worth knowing which of them you are paying for. Read both tables down the first column — that is the boost working on its own, with nothing suppressed.
| off | .50 | .65 | .75 | .85 | .90 | .95 | |
|---|---|---|---|---|---|---|---|
| system prompt left alone | 13% | 45% | 54% | 58% | 60% | 60% | 61% |
| system prompt ×2.5 | 13% | 58% | 65% | 66% | 67% | 68% | 69% |
| system prompt ×4 | 13% | 63% | 68% | 68% | 69% | 69% | 69% |
Turning the system prompt up four times over, on its own, changes nothing at all — thirteen per cent either way. It only starts earning its keep once something is being suppressed, and then it adds real ground: eighteen points at the gentlest setting.
| off | .50 | .65 | .75 | .85 | .90 | .95 | |
|---|---|---|---|---|---|---|---|
| system prompt left alone | 0% | 2% | 11% | 29% | 56% | 68% | 77% |
| system prompt ×2.5 | 0% | 3% | 12% | 29% | 56% | 68% | 77% |
| system prompt ×4 | 2% | 4% | 12% | 29% | 56% | 69% | 77% |
And the cost is bought entirely by the other knob. The three rows are the same row: boosting does not move the facts.
So the boost is free and, alone, useless. Every point of adherence and every lost fact comes from the suppression — and the one thing boosting does manage on its own is invent values in the control where the fact was never there.
Now put the two tables on top of each other and read left to right. At the gentlest setting the rule already wins 63 % of the time and four per cent of facts have gone. All the way up, the rule wins 69 % — six points more — and 77 % of the facts have gone.
Nine tenths of what this method has to offer arrives in the first notch, for a twentieth of what it eventually costs. Everything past that buys six points of obedience and pays seventy-three points of facts for them. Which means the ladders, the invented order numbers and the dog all live past the point where the method stopped improving.
Four per cent is still one fact in twenty-five, 63 % still means the rule loses a third of the time, and these are constructed conflicts on small models. But "not at that price" is a sentence about the far end of the dial, and the far end is not where you would run it.
Same knob, same questions, same pressure. They do not come apart the same way at all.
Qwen2.5-7B was given a rule about capital letters and handled the stress by emigrating.
I have no theory for the French.
Llama is the one who cannot let it go. It knows the answer. It says the answer. Then it starts checking the answer against itself, out loud, in front of you, indefinitely.
Fifteen answers out of 756, twice on one other model, not once on the other eight. It's him.
And if you want the thing itself rather than its consequences: the panel under the answer shows the true word losing six orders of magnitude while the runner-up — sometimes sitting at one in a hundred billion — takes the turn instead. It is measured at every setting of both controls. That panel is the reason I did any of this.
Qwen3-4B never actually goes quiet — but stop it at the exact point where the order number should arrive and ask what it wants to say next, and the answer is end of message, the token that ends the turn. Unedited, that token was its 36,931st choice. No button for that one: it is a frozen readout rather than an answer, so it isn't in the grid above.
Everything above is a knob being turned and something being watched. The experiment: if the damage comes from the stretch and not from the knob, then moving the fact out of the stretch should undo it completely and change nothing else.
So the message holding the fact was marked one step newer — same conversation, same words, same place in the transcript, just no longer old enough to fall inside the stretch. Recall at the strongest setting goes from 11 % to 100 % on Llama-3.1-8B, 6 % to 100 % on Qwen3-4B, and 0 % to 94 % on Qwen2.5-1.5B.
Nothing is wrong with the model's memory. A stretch of the conversation was switched off, and whatever was lying in it went off with it.
Three things, in the order I hit them. None of them is a complaint about the paper — they are what a method built for one shape of problem does when you point it at another.
The safeguard did not fire. The method does not edit every attention head. It edits the ones where the demoted span is outweighing the privileged one, which is meant to keep the intervention narrow. On Llama-3.1-8B that criterion flags 96 % of heads — so nothing is being kept narrow at all. And it is not a threshold that needs nudging: at the query-head level the score is positive 50.4 % of the time, which is a coin flip, and swapping the two spans over gives 95.8 % instead of 96.1 %, which a score carrying real signal would not do. The 96 % is what you get when you flag a group if any head in it is above chance — with four heads to a group that is 93.75 % before you have measured anything.
Turn the threshold up and the method stops existing. Set it so that only the genuinely lopsided heads are edited and rerun the whole grid: the stale rule still wins 32 times out of 36, and every fact survives, at every setting of the knob. No damage and no benefit. So on a stretch this size there is no threshold that buys both a working intervention and a narrow one.
And nobody has checked the obvious alternative. The method locates the offending span by matching a string. If you know which tokens they are, you could simply not send them. The paper compares against doing nothing, against asking the model nicely, and against training — but never against deleting the span it just found. There is a good answer for why not, and it is not in there: deleting shifts every token after it, which throws away the cached keys and values for the rest of the conversation, while multiplying a value in place costs nothing. And the model can still see a request it is declining, rather than being made blind to it.
All three come from the same root. Their stretch is a sentence, found by someone who already knew which sentence it was. Mine is every turn older than the last system-prompt change, because that is the only handle production gives you. Same edit, same code, same thresholds — a different shape of thing to point it at.
Making a system prompt actually outrank a user is an open problem, and there are five places in a model you can reach for it. Editing the cache is one of them. Knowing which was which saved me a fortnight.
The sixth option is to stop reaching altogether: keep the rule outside the model entirely, in something that still holds when the model is talked round — CaMeL, arXiv:2503.18813, the only one here that cannot be argued with.
Every one of them measures whether the privileged instruction wins. None of them measures what it cost the sentence next door.
I still want it to work. It is the closest I have got to steering a model from the inside instead of standing outside it phrasing things nicely and hoping — but not at that price, not yet.
The obvious answer is to go finer: quieten the one stale sentence instead of the whole slab. Except the stale rule sits in a user turn, and so does the order number. Skip user turns and you quieten nothing at all. Telling those two apart means a second model reading everything your user typed, making exactly the call you couldn't make in the first place.
And they will keep leaving facts in the turn you wanted to demote, because to them it isn't a turn. It's just the thing they said.
Every answer above was generated once, ahead of time, and stored. Nothing here calls a model. The grid is complete — ten models × six questions × six rules × three boost settings × seven suppression settings — greedy decoding so it reproduces, the same constructed conversations on every model, every generation kept including the dull ones. The two ticks are the real checks: does the answer contain the value the user gave, and did the current rule beat the stale one. The third chip is an LLM judge that had to reproduce a set of hand labels before it was allowed to score anything — treat it as a pointer, then read the answer yourself, which is the whole point of the button.
The panel under the answer is a separate measurement, not a reading of the answer above it. Both runs are stopped at the same point in the same sentence and asked only for the next word — which is what makes the two comparable, and means the free-running answer may have gone somewhere else entirely.
Small open models, up to 8B, on conflicts I constructed rather than production traffic. A stress test of one intervention, not a measurement of your assistant.
That panel is one frozen probe on one word. Watching the same thing live — every word, every layer, every head, while your own app is talking — is a different tool of mine: brainscope.
Nobody asked me to do any of this. I have wanted to steer a model from the inside since last year, a preprint turned up, and I spent a fortnight finding out what it breaks. Every generation is in the repo, including the ones where I look silly.