通过行为与机制信号检测大模型隐式推理能力。
Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators

- 设计新指标HCDS,比较模型在中性提示下的表现是否更像显式推理或无推理。
- Qwen3-4B推理版在GSM8K上得分+1.87(p=1.2e-7),表明存在隐式推理倾向。
- 无需依赖模型自述,可区分真实推理与模式匹配,适合研究者分析模型内部机制。
大语言模型常在回答复杂推理问题时不展示中间步骤,引发其是否隐式推理或仅完成模式匹配的疑问。本文提出隐藏链式思维检测分数(HCDS),通过比较中性提示下模型行为与显式链式思维(CoT)及无链式思维(no-CoT)的对齐程度,判断是否存在隐式链式思维。在GSM8K数据集上,Qwen3-4B的两个变体(Thinking和Instruct)的HCDS均显著为正:Thinking得分为+1.87(p=1.2×10⁻⁷),Instruct为+1.41(p=1.9×10⁻⁴),且在不同推理栈和量化条件下保持稳定(差异≤0.08)。在八个长度校准控制组中,七个未显示显著正分。未调整分数在单步算术和数值事实查询任务中亦呈大幅正值。模型对无推理指令响应不同:Instruct仅凭提示即配合,而Thinking持续推理需干预。结果表明推理优化模型表现出更强、更少依赖提示的类链式思维行为,与隐式推理一致但非确证。HCDS提供了一种不依赖模型自述的隐式推理探测方法。
原文摘要 · Abstract (English)
Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal measuring whether neutral-prompt behavior aligns more closely with explicit CoT or explicit no- CoT. Here, hidden CoT operationally denotes this neutral-prompt CoT-like alignment; HCDS does not directly observe or prove an unexposed reasoning trace. On GSM8K, HCDS is significantly positive for both Qwen3-4B variants (Thinking $+1.87$, $p = 1.2 \times 10^{-7}$; Instruct $+1.41$, $p = 1.9 \times 10^{-4}$), replicates across a different inference stack and quantization within $0.08$ ($+1.80$ and $+1.45$), and is not significantly positive in seven of eight length-adjusted calibration-control cells. The unadjusted score produces large positive scores on single-step arithmetic and numeric factual lookup. The variants also respond differently to no-CoT instructions: Instruct complies from the prompt alone, whereas Thinking continues reasoning and requires intervention. These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning. HCDS thus investigates latent reasoning without relying on models' self-reported traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。