September 2026

The Geometry of a Latent ThoughtCalling a hidden state a โ€œlatent thoughtโ€ is an assumption, borrowed from recent papers

Coconut-style models reason in continuous hidden states, but those latent thoughts often go unused. This post tests whether their geometry is the reason. Thoughts go unused when the task doesn't need them, and unused thoughts look collapsed. When the task does need them, the model uses them whatever their shape.

Background

CoconutHao et al. (2024), Coconut, abstract. lets a language model think without words. Instead of turning its hidden state into a word and reading that word back, it feeds the hidden state itself back in as the next input. Each fed-back vector is one "thought".

Chain of thought: through a word

language model hidden state pick one word "7" word embedding next input the whole state is rounded to one word

Coconut: straight back in

language model hidden state = the next input one "thought" <bot> thought loop <eot> no word, nothing rounded away

Training uses a curriculum. The model first learns ordinary step-by-step reasoning in words. Then, stage by stage, the first written step is swapped for a thought, then the first two, and so on, until only thoughts remain before the answer. At every stage the model is graded only on the words that are still written out and on the final answer. The thoughts themselves are never graded.

It's done this way because there is no "correct thought" to train against: nobody knows what the vector should contain. The curriculum works around that. At each stage the model only has to learn one new thought, and it has an immediate check on it, because the next written step has to follow from that thought. The thought ends up carrying whatever the model needs to continue. Without the curriculum it doesn't work: in the Coconut paper, training on thoughts directly scored no better than answering with no reasoning at allHao et al. (2024), Coconut, Table 1. Without the curriculum, ProsQA accuracy is 76.1%, against 76.7% with no reasoning (No-CoT). (Hao et al., Table 1). The idea comes from earlier work that gradually removed written steps so a model learns to reason internally (Deng et al., 2024Deng, Choi & Shieber (2024), abstract.).

Coconut's training curriculum, simplified from Figure 2 of Hao et al. (2024). Orange steps are written words and are graded; blue thoughts are never graded directly. The paper can use more than one thought per replaced step.

It is an elegant idea, but on standard benchmarks the thoughts often go unused. Models reach the same answers without them (Aswal et al.Aswal et al. (2026), Table 1. Coconut (row C) on graph-hopping, the ProsQA-style task: 98.0% with its thoughts, 97.8% with them removed.; Rizvi-Martel et al.Rizvi-Martel et al. (2026), conclusion, page 10.), and the Coconut paper's own pause-token variant nearly matches itHao et al. (2024), Coconut, Table 1. On ProsQA, โ€œpause as thoughtโ€ scores 96.6% against Coconutโ€™s 97.0%..

What does a thought look like?

I measured the vector Coconut would feed back in seven pretrained models (GPT-2, GPT-2-medium, Pythia-410M, SmolLM2-360M, Qwen2.5-0.5B and its Instruct version, and Qwen2.5-1.5B), on ordinary text: how long it is compared with the model's input word embeddings, and how similar it is to its own average.

Length and similarity of the fed-back vector in seven models
One row per pretrained model. Left: length of the vector Coconut would feed back, compared with the model's word embeddings. Right: how similar that vector is to its own average across tokens (1.0 = identical every time)
GPT-2 familyOther models

In every model it is 18โ€“581ร— longer than a word. In GPT-2, Coconut's usual base model, 94% of it sits in three dimensions and it points almost the same way for every token, a known quirk of GPT-2's representations (Timkey & van SchijndelTimkey & van Schijndel (2021), abstract.; Sun et al.Sun et al. (2024), abstract.). The Coconut paper says the fed-back states "are not too large in magnitude"Excerpt from the Coconut paper, Section 3: It is worth noting that the last hidden states have been processed by the final normalization layer, so they are not too large in magnitude.Hao et al. (2024), Coconut, Section 3, page 4 ยท arXiv 2412.06769v4. That holds relative to the un-normalized state, which the final normalization shrinks. Relative to the inputs the model reads, they are still huge, even in a trained Coconut model: I measured a median length of 172โ€“219 for the thoughts of a published ProsQA checkpoint, against about 3 for a typical GPT-2 word embedding, so roughly 60โ€“70ร— a word. To be fair, the Coconut authors do not intend thoughts to look like words: they state that the latent thought "is not intended to be mapped back to language space"Hao et al. (2024), Coconut, Section 3, page 4., and that their training objective does not push a thought to compress the written step it replacesHao et al. (2024), Coconut, Section 3, page 4.. The mismatch is a design choice, not an oversight. The question here is whether it costs anything.

Checking a known result on ProsQA

Several groups have reported that Coconut barely uses its thoughts on ProsQAHao et al. (2024), Coconut, Section 4.1, page 5., a two-choice logic puzzle introduced in the Coconut paper. The paper's own pause-token version nearly matches it (96.6% against 97.0%)Hao et al. (2024), Coconut, Table 1.; removing the thoughts barely changes answers (Aswal et al.Aswal et al. (2026), Table 1.; Rizvi-Martel et al.Rizvi-Martel et al. (2026), conclusion, page 10.); and the author of a published Coconut model found the same with corruption and transplant tests (bmarti44, 2026; discussion). Before building on this, I checked it on that same checkpoint (GPT-2 trained on ProsQA), removing its thoughts or replacing them with their average.

ProsQA accuracy with the thoughts removed or replaced
Each square is one of the 500 ProsQA test questions. Switch what the model is given and watch which squares change. Hover a square to see its question, and click it to read the whole problem. Chance is 50%, since every question has two options
98.0% correct0 of 500 answers differ from its own thoughts
answered correctlyanswered wronglychanged

Deleting the thoughts changed 1 answer in 500, and replacing them with their average changed none, which confirms that the model solves ProsQA from the question alone. The thoughts in this checkpoint also have the collapsed shape: roughly 60โ€“70ร— a word, 92% in three dimensions, and nearly identical across questions (similarity 0.994โ€“0.996 to their average), which is also why swapping in the average changes so little; deleting them is the stronger test. Here, unused thoughts and collapsed geometry go together. ProsQA therefore can't show whether the shape of a thought matters, because the task doesn't need thoughts at allHao et al. (2024), Coconut, Section 5.3, page 10. The โ€œtwo other synthetic tasksโ€ are ProntoQA and ProsQA., so the next step needs a different task.

A task that needs thinking

Picture five cups and a sequence of shuffles, such as "swap 1โ†”2 and 3โ†”4" or "rotate 3โ†’4โ†’5". Where does each cup end up? These even shuffles form a group called A5, and the natural solution updates the arrangement once per shuffle.

The Coconut authors already have one task where the thoughts do matter. On GSM8K, a set of grade-school math problems, Coconut scores 34.1% against 24.1% for the same curriculum with pause tokensHao et al. (2024), Coconut, Table 1 (GSM8k column)., and they explain the gap by GSM8K "placing higher demands on computational capability"Hao et al. (2024), Coconut, Section 5.3, page 10.. I use a synthetic task instead, for three reasons. Training Coconut on GSM8K takes far more compute than this study had. In a synthetic task every intermediate state is known, so I can check what each thought actually contains. And the difficulty can be set exactly by the number of steps, with theory saying how the need for thinking grows with it.

An example A5 problem
The model sees only the four moves. Reasoning in words, it writes the arrangement after each move; a Coconut model has to carry it inside its thoughts

    Theory says no fixed-depth transformer, pause tokens included, can do this for every length, while one reasoning step per shuffle can (Barrington, 1986Barrington (1986), STOC, Theorem 4, page 3. A5 is a non-solvable group. The end of โ€œnon-solvableโ€ is clipped in the original scan.; Merrill & SabharwalMerrill & Sabharwal (2025), abstract. Padding (pause) tokens keep a transformer within TCโฐ.). At short lengths, though, a model can still combine shuffles in parallel in about log n layers (Liu et al.Liu et al. (2023), abstract.), so the theory alone does not guarantee that thinking is needed here. In our tests, a 4-layer model with no thinking steps failed on problems with more than 4 shuffles.

    Does the shape matter when thoughts are needed?

    I trained 4-layer models from scratch, first to reason in words and then to replace the words with thoughts. I tried three versions of the thought: raw (as in Coconut), rescaled to word length, and centered then rescaled. With h the model's last hidden state, the vector fed back is:

    Rawh Rescaledr ยท h / โ€–hโ€– Centeredr ยท (h โˆ’ ฮผ) / โ€–h โˆ’ ฮผโ€–

    โ€–ยทโ€– is a vector's length; r is the average length of the model's word embeddings; ฮผ is a running average of hidden states, updated during training (ฮผ โ† 0.99 ฮผ + 0.01 ร— batch average) and then kept fixed.

    The test replaces every thought with the average thought. If accuracy survives that, the thoughts carry nothing.

    A5 accuracy with each model's own thoughts and with their average
    Accuracy on 8-shuffle A5 problems (1,000 problems each). Columns show the mean of 3 seeds; small dots are the individual seeds. Chance is about 1.7%
    With its own thoughtsThoughts replaced by their average
    The shape of each version's thoughts
    The vectors fed back on 8-shuffle problems, range across 3 seeds. Pretrained GPT-2 (from the first chart) shown for comparison
    VersionLength vs a wordSimilarity to average
    Pretrained GPT-2, for comparison82ร—0.99
    Coconut raw12โ€“17ร—0.34โ€“0.62
    Coconut rescaled1ร— (set)0.30โ€“0.34*
    Coconut centered1ร— (set)0.86โ€“0.96โ€ 

    * Measured on the hidden states before rescaling; rescaling changes the length, not the direction. โ€  Measured on the hidden states before centering. The average is subtracted before feeding back, so what the model receives is spread out, even though its hidden states became more alike.

    All three versions solved 8-shuffle problems on every seed, and the swap sent every one to chance. Their thoughts decode to the correct intermediate arrangement 93โ€“100% of the time. The shape made no difference. Even the raw thoughts were not collapsed here: they were long, but they differed from problem to problem, unlike pretrained GPT-2's. One caveat: the average thought is itself an unusual input.

    Does collapsed geometry block a pretrained model?

    The tiny models never had GPT-2's collapsed geometry, so I fine-tuned pretrained GPT-2 on the same task. When Coconut training began, its fed-back vector was 26ร— a word, 90% in three dimensions, and nearly identical across problems (similarity 0.96). Coconut learned the task anyway, and while doing so it reshaped its own thoughts. By the end of training they were 4.5ร— a word, 27% in three dimensions, and far less alike (similarity 0.36).

    A pretrained model learning to use its thoughts
    training step0
    accuracy2%
    similarity to average0.96
    length vs a word26ร—
    collapsed ยท thoughts unused
    drag to scrub through training

    A schematic, not a projection of real thoughts: the fan of arrows is simulated, and only its spread and length are set from measurements. They follow pretrained GPT-2's measured similarity to the average (0.96 โ†’ 0.36) and length (26ร— โ†’ 4.5ร— a word, compressed for display) during Coconut training on the A5 task. One training run (a single seed).

    Accuracy and thought similarity during training
    Pretrained GPT-2 fine-tuned on A5 (8 shuffles), one seed, measured every 1,000 training steps on 500 problems. The two panels share the time axis.

    Accuracy rose from chance to 99%. Within the first 2,000 steps, the thoughts' similarity to their average fell from 0.96 to 0.27, and their length from 26ร— to about 4ร— a word. Rescaled and centered versions did no better, and all three fall to chance under the average swap.

    Collapsed geometry didn't stop the model from learning to use its thoughts.

    Conclusion

    What we found

    • Thoughts go unused when the task doesn't need them, and unused thoughts look collapsed.
    • When the task needs them, the model uses them whatever their shape, even starting from GPT-2's collapsed geometry.
    • Hand-fixing the shape (rescaling or centering) adds nothing.

    How it fits earlier work

    • Earlier work links collapse to disuse (Aswal et al.Aswal et al. (2026), introduction, page 2.; SIM-CoTWei et al. (2025), SIM-CoT, abstract.). What this adds is that collapse did not prevent use.
    • It fits the finding that recycled thoughts can hurt generalization to longer problems and make the model overconfident, and that what fills the thought slots matters less than how the model is trained (bmarti44, 2026).

    Limits

    • The pretrained result is one seed.
    • The models are small (up to 124M trained) and the task is a toy.
    • Final accuracy is at ceiling, so effects on learning speed are untested.

    References

    1. Hao et al. Training Large Language Models to Reason in a Continuous Latent Space (Coconut). arXiv 2412.06769
    2. Aswal et al. Observable Patterns Are Not Explanations. arXiv 2606.12689
    3. Rizvi-Martel et al. Illusion of Superposition? arXiv 2604.06374
    4. Wei et al. SIM-CoT. arXiv 2509.20317
    5. Deng, Choi & Shieber. From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step. arXiv 2405.14838
    6. bmarti44. The Curriculum Is the Mechanism (manuscript, 2026). PDF ยท checkpoints ยท r/MachineLearning discussion
    7. Timkey & van Schijndel. All Bark and No Bite: Rogue Dimensions in Transformer Language Models. arXiv 2109.04404
    8. Sun et al. Massive Activations in Large Language Models. arXiv 2402.17762
    9. Barrington. Bounded-width polynomial-size branching programs recognize exactly those languages in NCยน. STOC 1986 (journal version: Journal of Computer and System Sciences, 1989). ACM
    10. Merrill & Sabharwal. Exact Expressive Power of Transformers with Padding. arXiv 2505.18948
    11. Liu et al. Transformers Learn Shortcuts to Automata. arXiv 2210.10749