arXiv:2607.26845cs.LG2026-07

研究大模型在不确定下的思考行为,发现思考能提升决策质量但不增强信息探索。

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

论文配图:Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models
图 1 · 摘自论文原文
  • 通过双臂老虎机实验区分思考对行动与信息寻求的影响
  • 思考使模型更依赖证据、减少随机选择,但未增加探索倾向
  • 思考长度和信心反馈体现元认知控制与监控,适合关注模型决策机制的研究者

推理时的思考能提升大语言模型性能,但整体表现无法揭示模型是更有效地利用已有证据,还是主动寻求信息以改进未来决策。我们通过测量动作偏好、思考长度和报告信心,在相同不确定性条件下区分这两种响应。十种开源模型在思维模式与非思维模式下完成匹配的多轮双臂老虎机任务。认知模型将价值引导的行为与独立于不确定性的随机噪声分离,识别出两类探索信号:类似上置信界(UCB)的对未知选项偏好,以及类似汤普森采样(Thompson)的随总不确定性增加的选择变异性。平均而言,思考增强了价值引导行为,减少了随机噪声,但未引发类似UCB的探索或强化类似Thompson的探索。在非动作阶段,信息不平衡历史条件(观察次数多于平衡条件)导致更长的思考长度。报告信心对决策难度更敏感,且与所选任务证据关联更强。这些模式分别支持元认知控制和元认知监控的解释,但未证实具体过程。解码器超参(尤其是温度)影响噪声和思考长度,但无法复现联合输出模式。在该受控决策设置中,思考提升了模型基于当前证据的行动能力,但两种测量均未支持向更具信息探索倾向策略的转变。

原文摘要 · Abstract (English)

Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking and non-thinking modes. A cognitive model separated value-guided action and uncertainty-independent choice noise from two behavioral signatures of exploration: a UCB-like preference for the less-known arm and Thompson-like choice variability that increases with total uncertainty. On average, thinking strengthened value-guided action and reduced uncertainty-independent choice noise, without producing UCB-like exploration or strengthening Thompson-like exploration. Outside action, the information-imbalanced history condition, which also displayed more observations than the matched balanced condition, was associated with greater thinking length. Reported confidence became more sensitive to decision difficulty and more strongly associated with chosen task evidence. We interpret these thinking-length and reported-confidence patterns as consistent with metacognitive control and metacognitive monitoring, respectively, without establishing either process. Decoder sweeps, especially temperature, altered choice noise and thinking length but did not reproduce the joint cross-output pattern. In this controlled decision setting, thinking improved how models acted on current evidence, while neither measured signature supported a shift toward a more information-seeking policy.

大模型推理元认知决策机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。