arXiv:2602.08693cs.LG2026-02被引 2

链式思维让大模型决策更像人类,尤其在推理阶段。

Reasoning aligns language models to human cognition

  • 设计新任务分离采样与推理,量化模型行为
  • 长链条推理显著提升推理能力,使信念变化趋近人类
  • 揭示模型与人类在信息获取上的根本差异

语言模型在不确定性下的决策是否类人?我们引入一项主动概率推理任务,清晰分离主动采样(获取证据)与推理(整合证据做决策)。对比人类与多种主流大语言模型与近最优参考策略的表现发现:延长推理过程是性能提升的关键,大幅改善推理表现,使信念演化轨迹高度类人,但对主动采样仅带来小幅提升。通过拟合一个包含记忆、策略、选择偏差和遮蔽感知四个可解释潜变量的机制模型,将人类与模型置于共享的低维认知空间中,再现跨主体的行为特征,并揭示链式思维如何推动语言模型向人类式的证据累积与信念-决策映射靠拢,强化了推理层面的对齐,却仍存在信息获取上的持久差距。

原文摘要 · Abstract (English)

Do language models make decisions under uncertainty like humans do, and what role does chain-of-thought (CoT) reasoning play in the underlying decision process? We introduce an active probabilistic reasoning task that cleanly separates sampling (actively acquiring evidence) from inference (integrating evidence toward a decision). Benchmarking humans and a broad set of contemporary large language models against near-optimal reference policies reveals a consistent pattern: extended reasoning is the key determinant of strong performance, driving large gains in inference and producing belief trajectories that become strikingly human-like, while yielding only modest improvements in active sampling. To explain these differences, we fit a mechanistic model that captures systematic deviations from optimal behavior via four interpretable latent variables: memory, strategy, choice bias, and occlusion awareness. This model places humans and models in a shared low-dimensional cognitive space, reproduces behavioral signatures across agents, and shows how chain-of-thought shifts language models toward human-like regimes of evidence accumulation and belief-to-choice mapping, tightening alignment in inference while leaving a persistent gap in information acquisition.

认知对齐链式思维人类决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。