Coconut-style models reason in continuous hidden states, but those latent thoughts often go unused. This post tests whether their geometry is the reason. Thoughts go unused when the task doesn't need them, and unused thoughts look collapsed. When the task does need them, the model uses them whatever their shape.
CoconutHao et al. (2024), Coconut, abstract. lets a language model think without words. Instead of turning its hidden state into a word and reading that word back, it feeds the hidden state itself back in as the next input. Each fed-back vector is one "thought".
Training uses a curriculum. The model first learns ordinary step-by-step reasoning in words. Then, stage by stage, the first written step is swapped for a thought, then the first two, and so on, until only thoughts remain before the answer. At every stage the model is graded only on the words that are still written out and on the final answer. The thoughts themselves are never graded.
It's done this way because there is no "correct thought" to train against: nobody knows what the vector should contain. The curriculum works around that. At each stage the model only has to learn one new thought, and it has an immediate check on it, because the next written step has to follow from that thought. The thought ends up carrying whatever the model needs to continue. Without the curriculum it doesn't work: in the Coconut paper, training on thoughts directly scored no better than answering with no reasoning at allHao et al. (2024), Coconut, Table 1. Without the curriculum, ProsQA accuracy is 76.1%, against 76.7% with no reasoning (No-CoT). (Hao et al., Table 1). The idea comes from earlier work that gradually removed written steps so a model learns to reason internally (Deng et al., 2024Deng, Choi & Shieber (2024), abstract.).
Coconut's training curriculum, simplified from Figure 2 of Hao et al. (2024). Orange steps are written words and are graded; blue thoughts are never graded directly. The paper can use more than one thought per replaced step.
It is an elegant idea, but on standard benchmarks the thoughts often go unused. Models reach the same answers without them (Aswal et al.Aswal et al. (2026), Table 1. Coconut (row C) on graph-hopping, the ProsQA-style task: 98.0% with its thoughts, 97.8% with them removed.; Rizvi-Martel et al.Rizvi-Martel et al. (2026), conclusion, page 10.), and the Coconut paper's own pause-token variant nearly matches itHao et al. (2024), Coconut, Table 1. On ProsQA, โpause as thoughtโ scores 96.6% against Coconutโs 97.0%..
I measured the vector Coconut would feed back in seven pretrained models (GPT-2, GPT-2-medium, Pythia-410M, SmolLM2-360M, Qwen2.5-0.5B and its Instruct version, and Qwen2.5-1.5B), on ordinary text: how long it is compared with the model's input word embeddings, and how similar it is to its own average.
In every model it is 18โ581ร longer than a word. In GPT-2, Coconut's usual base model, 94% of it sits in three dimensions and it points almost the same way for every token, a known quirk of GPT-2's representations (Timkey & van SchijndelTimkey & van Schijndel (2021), abstract.; Sun et al.Sun et al. (2024), abstract.). The Coconut paper says the fed-back states "are not too large in magnitude"Hao et al. (2024), Coconut, Section 3, page 4 ยท arXiv 2412.06769v4. That holds relative to the un-normalized state, which the final normalization shrinks. Relative to the inputs the model reads, they are still huge, even in a trained Coconut model: I measured a median length of 172โ219 for the thoughts of a published ProsQA checkpoint, against about 3 for a typical GPT-2 word embedding, so roughly 60โ70ร a word. To be fair, the Coconut authors do not intend thoughts to look like words: they state that the latent thought "is not intended to be mapped back to language space"Hao et al. (2024), Coconut, Section 3, page 4., and that their training objective does not push a thought to compress the written step it replacesHao et al. (2024), Coconut, Section 3, page 4.. The mismatch is a design choice, not an oversight. The question here is whether it costs anything.
Several groups have reported that Coconut barely uses its thoughts on ProsQAHao et al. (2024), Coconut, Section 4.1, page 5., a two-choice logic puzzle introduced in the Coconut paper. The paper's own pause-token version nearly matches it (96.6% against 97.0%)Hao et al. (2024), Coconut, Table 1.; removing the thoughts barely changes answers (Aswal et al.Aswal et al. (2026), Table 1.; Rizvi-Martel et al.Rizvi-Martel et al. (2026), conclusion, page 10.); and the author of a published Coconut model found the same with corruption and transplant tests (bmarti44, 2026; discussion). Before building on this, I checked it on that same checkpoint (GPT-2 trained on ProsQA), removing its thoughts or replacing them with their average.
Deleting the thoughts changed 1 answer in 500, and replacing them with their average changed none, which confirms that the model solves ProsQA from the question alone. The thoughts in this checkpoint also have the collapsed shape: roughly 60โ70ร a word, 92% in three dimensions, and nearly identical across questions (similarity 0.994โ0.996 to their average), which is also why swapping in the average changes so little; deleting them is the stronger test. Here, unused thoughts and collapsed geometry go together. ProsQA therefore can't show whether the shape of a thought matters, because the task doesn't need thoughts at allHao et al. (2024), Coconut, Section 5.3, page 10. The โtwo other synthetic tasksโ are ProntoQA and ProsQA., so the next step needs a different task.
Picture five cups and a sequence of shuffles, such as "swap 1โ2 and 3โ4" or "rotate 3โ4โ5". Where does each cup end up? These even shuffles form a group called A5, and the natural solution updates the arrangement once per shuffle.
The Coconut authors already have one task where the thoughts do matter. On GSM8K, a set of grade-school math problems, Coconut scores 34.1% against 24.1% for the same curriculum with pause tokensHao et al. (2024), Coconut, Table 1 (GSM8k column)., and they explain the gap by GSM8K "placing higher demands on computational capability"Hao et al. (2024), Coconut, Section 5.3, page 10.. I use a synthetic task instead, for three reasons. Training Coconut on GSM8K takes far more compute than this study had. In a synthetic task every intermediate state is known, so I can check what each thought actually contains. And the difficulty can be set exactly by the number of steps, with theory saying how the need for thinking grows with it.
Theory says no fixed-depth transformer, pause tokens included, can do this for every length, while one reasoning step per shuffle can (Barrington, 1986Barrington (1986), STOC, Theorem 4, page 3. A5 is a non-solvable group. The end of โnon-solvableโ is clipped in the original scan.; Merrill & SabharwalMerrill & Sabharwal (2025), abstract. Padding (pause) tokens keep a transformer within TCโฐ.). At short lengths, though, a model can still combine shuffles in parallel in about log n layers (Liu et al.Liu et al. (2023), abstract.), so the theory alone does not guarantee that thinking is needed here. In our tests, a 4-layer model with no thinking steps failed on problems with more than 4 shuffles.
I trained 4-layer models from scratch, first to reason in words and then to replace the words with thoughts. I tried three versions of the thought: raw (as in Coconut), rescaled to word length, and centered then rescaled. With h the model's last hidden state, the vector fed back is:
โยทโ is a vector's length; r is the average length of the model's word embeddings; ฮผ is a running average of hidden states, updated during training (ฮผ โ 0.99 ฮผ + 0.01 ร batch average) and then kept fixed.
The test replaces every thought with the average thought. If accuracy survives that, the thoughts carry nothing.
| Version | Length vs a word | Similarity to average |
|---|---|---|
| Pretrained GPT-2, for comparison | 82ร | 0.99 |
| Coconut raw | 12โ17ร | 0.34โ0.62 |
| Coconut rescaled | 1ร (set) | 0.30โ0.34* |
| Coconut centered | 1ร (set) | 0.86โ0.96โ |
* Measured on the hidden states before rescaling; rescaling changes the length, not the direction. โ Measured on the hidden states before centering. The average is subtracted before feeding back, so what the model receives is spread out, even though its hidden states became more alike.
All three versions solved 8-shuffle problems on every seed, and the swap sent every one to chance. Their thoughts decode to the correct intermediate arrangement 93โ100% of the time. The shape made no difference. Even the raw thoughts were not collapsed here: they were long, but they differed from problem to problem, unlike pretrained GPT-2's. One caveat: the average thought is itself an unusual input.
The tiny models never had GPT-2's collapsed geometry, so I fine-tuned pretrained GPT-2 on the same task. When Coconut training began, its fed-back vector was 26ร a word, 90% in three dimensions, and nearly identical across problems (similarity 0.96). Coconut learned the task anyway, and while doing so it reshaped its own thoughts. By the end of training they were 4.5ร a word, 27% in three dimensions, and far less alike (similarity 0.36).
A schematic, not a projection of real thoughts: the fan of arrows is simulated, and only its spread and length are set from measurements. They follow pretrained GPT-2's measured similarity to the average (0.96 โ 0.36) and length (26ร โ 4.5ร a word, compressed for display) during Coconut training on the A5 task. One training run (a single seed).
Accuracy rose from chance to 99%. Within the first 2,000 steps, the thoughts' similarity to their average fell from 0.96 to 0.27, and their length from 26ร to about 4ร a word. Rescaled and centered versions did no better, and all three fall to chance under the average swap.
Collapsed geometry didn't stop the model from learning to use its thoughts.
What we found
How it fits earlier work
Limits