Does Making an AI "Think Longer" Also Make It Hallucinate More?
Why making AI models think longer causes visual attention drift, confabulation loops, and the paradox of cognitive overthinking.
Published: Sep 30, 2026
4 mins read
👁️ 34 Unique Views
Modern AI strategy relies on a simple premise: more compute equals better performance.They generate thousands of internal reasoning tokens before answering. This works brilliantly for math, coding, and formal logic. However, Vision-Language Models (VLMs) present a major drawback. Extended reasoning chains often increase hallucinations and erode visual accuracy.
"More Thinking, Less Seeing?"
A landmark paper at NeurIPS 2025 confirmed this vulnerability (More Thinking, Less Seeing?, Liu et al., arXiv:2505.21523). Researchers introduced the RH-Bench benchmark and the RH-AUC metric to test how reasoning length affects visual accuracy.
The results were clear. Extra thinking helped multi-step visual math. But on basic perceptual tasks, longer reasoning led to a steady rise in hallucinated objects, fake colors, and wrong spatial layouts.
Attention Drift: How Internal Tokens Distract the Model
The main issue is attention drift. Transformers distribute attention across all sequence tokens. Visual inputs use fixed patch embeddings, usually between 576 and 1,024 tokens.
In short responses, attention stays pinned to those visual patches. But as the reasoning chain grows into thousands of text tokens, those text tokens overpower the visual input. The model stops looking at the image and relies on its own generated text.
Four Failure Modes of Overthinking
Mechanistic studies show that overthinking triggers four specific failure modes:
Rationalization Loops: A minor visual mistake in the first 100 tokens gets locked in. The model builds an elaborate narrative to justify its initial error instead of correcting it.
Dominance of Text Patterns: As visual attention fades, the model defaults to standard internet text patterns. Ask about a "kitchen counter," and it will list blenders, toasters, and spice racks that aren't in the photo.
Answer Abandonment: Research shows models often spot the correct visual answer within 150 tokens. After 1,500 tokens of overthinking, they second-guess themselves and pick a wrong answer.
Phantom Detail Fabrication: Forced deep deliberation makes models invent non-existent details. They add clock hands or architectural features just to fill the compute budget.
Where Test-Time Compute Helps vs. Hurts
Reasoning length depends on the task. Multi-step visual math platforms like MathVista benefit from extra compute because symbolic logic matters most.
On perceptual tasks—like object counting, medical radiography, or hazard detection—extra compute hurts performance. Direct, short-response models outperform over-deliberating models by 15% to 28% in factual precision.
Fixing the Problem
Researchers are testing three architectural constraints to keep reasoning grounded:
Gated Perception-Reasoning Optimization (GPRO): Controllers route simple visual questions to fast System-1 processing, saving deep System-2 deliberation for symbolic logic.
Forced Visual Re-Grounding: Checkpoints every 256 tokens force the model to re-examine visual patch embeddings.
Early Stopping: Inference stops as soon as the model reaches semantic consensus, rather than running through full token limits.
More Reasoning Is Not Better Reasoning
More compute does not automatically mean higher intelligence. Human cognition balances sensory perception with reflection. A driver does not deliberate for two minutes over a red light.
In multimodal AI, unanchored thinking causes errors. True intelligence requires knowing when to think deeply, and when to stop and look at the evidence. In multimodal AI, unanchored thinking causes errors. True intelligence requires knowing when to think deeply, and when to stop and look at the evidence.