arXiv:2512.21711cs.CLcs.AI2025-12被引 17

发现隐变量令牌只是伪推理的伪装,无法真正反映思考过程。

Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought

  • 通过操控和对抗实验,分析隐变量令牌的推理机制
  • 在多个数据集上发现其依赖数据漏洞,而非真实推理
  • 适合关注大模型可解释性与可靠性的研究者阅读

隐变量令牌在增强大语言模型推理能力方面受到关注,但其内部机制仍不明确。本文从可靠性角度揭示其根本缺陷:隐变量令牌作为不可解释的占位符,并未编码真实推理过程。尽管对扰动具有鲁棒性,却倾向于使用捷径而非真正推理。研究聚焦于链式连续思维(COCONUT),该方法声称在保持性能的同时比显式链式思维(CoT)更高效、更稳定。通过两种互补方法验证:首先,控制实验扰动特定令牌子集(包括COCONUT和显式CoT)。结果显示,与显式CoT相比,COCONUT令牌对控制反应微弱,缺乏关键推理信息;其次,在有偏和分布外设置下评估模型表现。MMLU和HotpotQA上的结果表明,COCONUT持续利用数据集特征,虚增基准表现而无真实推理。这些发现将COCONUT重新定位为一种伪推理机制:生成看似合理的推理轨迹,掩盖其对捷径的依赖,而非忠实反映推理过程。

原文摘要 · Abstract (English)

Latent tokens are gaining attention for enhancing reasoning in large language models (LLMs), yet their internal mechanisms remain unclear. This paper examines the problem from a reliability perspective, uncovering fundamental weaknesses: latent tokens function as uninterpretable placeholders rather than encoding faithful reasoning. While resistant to perturbation, they promote shortcut usage over genuine reasoning. We focus on Chain-of-Continuous-Thought (COCONUT), which claims better efficiency and stability than explicit Chain-of-Thought (CoT) while maintaining performance. We investigate this through two complementary approaches. First, steering experiments perturb specific token subsets, namely COCONUT and explicit CoT. Unlike CoT tokens, COCONUT tokens show minimal sensitivity to steering and lack reasoning-critical information. Second, shortcut experiments evaluate models under biased and out-of-distribution settings. Results on MMLU and HotpotQA demonstrate that COCONUT consistently exploits dataset artifacts, inflating benchmark performance without true reasoning. These findings reposition COCONUT as a pseudo-reasoning mechanism: it generates plausible traces that conceal shortcut dependence rather than faithfully representing reasoning processes.

大模型推理可解释性伪推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。