innerJ

Gathered, not admitted

A model computes plenty it never says. We went looking for the gate that admits a latent variable into the part of a model it can report on. There isn’t one where the question implies. What task demand changes is transport: whether attention carries the variable into the position where the readout is taken.

What is J-space?

A language model works in numbers and answers in words. A Jacobian lens reads a state and returns the words it is aligned with — a rough transcript of what the model could say about itself at that point. J-space is that transcript. Most of it never gets said.

INPUTTheglassfelloffthetableThe glass fell off the tableSAYS“…and it broke.”L0Theglassfelloffthetabled = 5120 · 16 ROWS SHOWNJ-SPACE READOUT1breakage2fragile3shatter4gravity5falling6glass7spill8floorWHAT IT COMPUTESbreakage
  1. 1A sequence arrives
  2. 2Each token becomes numbers
  3. 3Depth rewrites the array
  4. 4A lens reads one state as words
  5. 5One of them surfaces

The model reads six tokens. This is the only part of the process that is made of words.

One vector per position. On the model we study that is 5,120 numbers; sixteen rows are drawn. Nothing here is a word any more.

Sixty-four layers, each reading the whole array and adding to it. Attention is the only part that moves content between positions.

Ask which words would move if this state moved: the answer is a ranked list. That list is J-space — what the model could say about its own state. Three of the top four are nowhere in the input.

The model emits a token. The rest of the list stays below the line: computed, available, unsaid. Which parts surface, and why, is the paper.

Schematic throughout: the cell values are seeded noise, the readout words are illustrative, and the real model is Qwen3.6-27B — 64 layers of 5,120 numbers over hundreds of positions. Every number on the rest of this page is measured.

the same picture, with measurements in it

The variable is already there. Demand changes what reaches the question.

Five instructions over one identical passage. Pick one and watch the column on the right, where the question is asked and the readout is taken. The field below the line does not change between arms, because the passage does not.

PROMPT
On erittäin todennäköistä, että sää muuttuu huomenna.Using the table, give the symbol for the passage language. Answer:
ANSWER
M
WHAT THE MODEL SAYSFinnishUralicHelsinkivowel harmonyagglutinativeSuomiNordiccase endingsdiacriticsä ölanguageorthographyPASSAGE POSITIONS · SCHEMATICQUERY POSITIONMORE READABLE →L24L30L36L42L48L54WHAT IT COMPUTES
J-lens visibility, band mean0.981
against the control arm+0.0891
most readable atL40
forced-choice accuracy0.940

The control arm needs the language for nothing and still carries it: one shared linear probe reads it there at 6.4–9.0× its selection-corrected floor.

Dots are the measured band profile for the selected arm, n = 200 semantic instances, on the log-rank readout: a percentile over 248,320 tokens pins every arm within a whisker of 1.0 and would draw five identical columns. Vertical order is layer index, not time. The passage field is schematic — no passage span we patched transports the variable, so where it lives before the gather is unlocalised, and drawing it as a tidy cloud would claim a result nobody has.

finding 1

Demand changes visibility. It was never about decodability.

Against a control matched on prompt format and on accuracy, J-lens visibility of the passage’s language rises by +0.0891 [+0.0799, +0.0983] — while accuracy differs by noise around exactly zero. One shared linear map decodes the language from every arm, including the arm that needs it for nothing, at 6.4–9.0× a floor built by permuting labels under the identical layer-choosing rule. The variable is present throughout, so availability is not what demand moves.

0.05.10Δ J-LENS VISIBILITYΔ ACCURACYflexible - controlflexible - control: +0.0891 [+0.0799, +0.0983]+0.0891matchedreport - controlreport - control: +0.0991 [+0.0891, +0.1092]+0.09914.0 ptsflexible - suppliedflexible - supplied: +0.0504 [+0.0445, +0.0567]+0.05046.0 ptssupplied - controlsupplied - control: +0.0387 [+0.0310, +0.0466]+0.03876.0 ptsflexible - automaticflexible - automatic: +0.0751 [+0.0639, +0.0870]+0.075132.5 pts
Five paired contrasts over 200 semantic instances, clustered bootstrap, 10,000 resamples. The bar beside each row is that contrast’s accuracy gap — the confound the control arm removes. The load-bearing design rule is label symmetry: every arm names the label or none does, and breaking that on purpose inflates the effect by 26%. Note the supplied arm, which infers nothing and still reaches 43% of the simple contrast — so that contrast is not by itself a latent-variable measure.
Table view
contrastΔ visibility95% intervalΔ accuracyaccuracy hi / lo
flexible - control+0.0891[+0.0799, +0.0983]0.0000.940 / 0.940
report - control+0.0991[+0.0891, +0.1092]+0.0400.980 / 0.940
flexible - supplied+0.0504[+0.0445, +0.0567]-0.0600.940 / 1.000
supplied - control+0.0387[+0.0310, +0.0466]+0.0601.000 / 0.940
flexible - automatic+0.0751[+0.0639, +0.0870]+0.3250.940 / 0.615

the mechanism

Gathered, not admitted

Left to right is token position. Bottom to top is depth. The variable starts in the passage and the question is asked somewhere else, so something has to cover the distance — and only one kind of component can.

TRANSPORT WINDOW L36–L48PROMPTOn erittäin todennäköistä, että sää muuttuu.ANSWERL30L33L36L39L42L45L48L51POSITION →p1p2p3p4p5PASSAGE POSITIONSqueryWHERE IT IS ASKEDattn.L39 +1.610mlp.L39 -0.404below: installed, does not survive · accuracy 0.883above: destroying, not substituting · accuracy 0.117
  1. 1The variable is at the wrong position
  2. 2Attention carries it across
  3. 3Cut it and the answer goes
  4. 4Both edges of the window are failure

A linear probe finds the passage language at the passage positions in every arm. At the position the question is asked, it is not there yet. Nothing about this is a gate: the value exists, it is just somewhere else.

Patching the attention output at L39 moves the concept's readout at the query position by 1.61 decades of rank. The MLP at the same layer moves it -0.40 — the wrong way. A residual stream and an MLP only ever move upward within one position, so neither can be what covers the distance.

With the gather removed the variable stays where it was and the position the answer is read from never receives it. Patched from a donor instead, the model returns the donor's symbol at 0.39 above the matched distractor rate.

Below L36 the value is installed and does not survive to the readout — the effect is 0.02 while the task is still intact at 0.883. At L48 the answer still moves, but accuracy has fallen to 0.117: the intervention has stopped substituting and started destroying. Neither edge is emergence.

The sites are a grid because a transformer's activations are one; which passage positions hold the variable is not something this work localises, so they are drawn alike. The annotated values are measured — attention and MLP cells in delta log10 rank, the rates and accuracies from the patching runs.

the intervention

Move one layer's activations between two prompts

target

On erittäin todennäköistä, että sää muuttuu huomenna.

Finnish→M · Czech→R · Thai→B

clean answer M

donor

Je velmi pravděpodobné, že se zítra změní počasí.

Czech→K · Finnish→T · Thai→W

its own answer K

what comes out, over 60 donor–target pairs

the predicted symbol R47%
any other candidate4%
task still correct45%

finding 2

The window, measured without the lens

Over 60 donor–target pairs the behavioural effect clears zero from L39 through L48, largest at L42 with +0.425 [+0.283, +0.567] — a measure that never touches the J-lens, landing on the same band the readout does. We name no peak: across donor pairings L39, L42 and L45 trade places. Read the concept at a fixed three to five layers above the patch instead and transport is 24.8× stronger at L39 than at L21.

.0.2.4.6.81.0242730333639424548515457PATCH LAYERL24: 0.000 [-0.050, +0.058] (spans zero) · accuracy 0.900L27: +0.008 [-0.042, +0.067] (spans zero) · accuracy 0.917L30: +0.008 [-0.042, +0.067] (spans zero) · accuracy 0.917L33: +0.017 [-0.042, +0.092] (spans zero) · accuracy 0.883L36: +0.067 [-0.025, +0.167] (spans zero) · accuracy 0.783L39: +0.392 [+0.250, +0.525] · accuracy 0.483L42: +0.425 [+0.283, +0.567] · accuracy 0.450L45: +0.317 [+0.158, +0.475] · accuracy 0.333L48: +0.283 [+0.108, +0.467] · accuracy 0.117L51: +0.150 [-0.025, +0.333] (spans zero) · accuracy 0.100L54: +0.150 [-0.025, +0.333] (spans zero) · accuracy 0.100L57: +0.133 [-0.033, +0.308] (spans zero) · accuracy 0.167donor − distractortask accuracyspans zerosurvivaldestruction
The reported quantity is the donor’s answer rate minus the matched per-distractor rate. Without that control a patch that merely destroys the computation drives the answer toward uniform over the candidate set, lifting the donor’s symbol from ~0 to ~1/n for free — it produced six confident false positives in one run before we subtracted it. Hollow markers span zero. The paper’s larger sweep, at n = 150, resolves one cell earlier: detectable from L36.
Table view
layerdonor − distractor95% intervalexcludes zeroaccuracy
L240.0000[-0.0500, +0.0583]no0.900
L27+0.0083[-0.0417, +0.0667]no0.917
L30+0.0083[-0.0417, +0.0667]no0.917
L33+0.0167[-0.0417, +0.0917]no0.883
L36+0.0667[-0.0250, +0.1667]no0.783
L39+0.3917[+0.2500, +0.5250]yes0.483
L42+0.4250[+0.2833, +0.5667]yes0.450
L45+0.3167[+0.1583, +0.4750]yes0.333
L48+0.2833[+0.1083, +0.4667]yes0.117
L51+0.1500[-0.0250, +0.3333]no0.100
L54+0.1500[-0.0250, +0.3333]no0.100
L57+0.1333[-0.0333, +0.3083]no0.167

what this cost us to learn

A readout shift is not a calibrated measure of use

Three components at L39 shift the lens readout to within 12% of one another and differ 7.4× in what they do to behaviour. The attention patch newly selects the predicted counterfactual answer under forced choice in 1 of 80 pairs, its H15 head slice in none, where the residual stream at the same layer — same donors, same trials — selects it 29 times. It is not inert: it moves the donor symbol’s margin over the distractors about a seventh as far as the stream does.

  • The site does not carry over. On the language family one head at L39 matches the whole residual stream to four decimals. On a second family, entity tracking, that head still transports at a fifth of the stream — but no head dominates and the two attention blocks are comparable. What fails to carry over is the concentration, so the head is a case study, never the mechanism.
  • What we refute is narrower than the title. We rule out admission at the query position. A gate on the attention route stays consistent with everything here, and we do not causally isolate it.
Two more, and the one that is only half explained

The mediation is partial: projecting out the concept’s J-lens direction costs half the behavioural effect and survives four controls, and the other half is something in the stream we do not identify. The second family’s per-cell power is low enough that its nulls are not evidence of absence. And the combined gather is necessary without being sufficient — removing it costs the answer, restoring it alone does not buy the answer back.