arXiv:2606.11016cs.AI2026-06

LLM做决策时有隐性优先级,但解释常不准确。

Superficial Beliefs in LLM Decision-Making

论文配图:Superficial Beliefs in LLM Decision-Making
图 1 · 摘自论文原文
  • 用属性渐变的虚拟选项测试模型选择行为
  • 模型行为可预测,但自述理由仅部分匹配真实驱动因素
  • 揭示了模型存在表面化信念,适合研究决策机制者阅读

我们探究大型语言模型在二选一情境中,其选择是仅模仿理由,还是反映系统性的内在决策结构。通过合成的二元决策场景,模型需在具有渐变属性的选项间做出选择。比较模型声称最重要的属性与基于历史选择拟合的行为模型所推断出的主导属性,发现行为模型能有效预测未见选择,表明模型行为与可见属性系统相关而非随机。然而,直接自述和评分裁判仅部分恢复行为模型推断出的驱动因素。结果表明,模型行为虽具结构性,但其明确理由与真实驱动因素匹配度有限。该模式在提示顺序、采样扰动、不同行为模型及结构变化下均保持一致,支持‘表面信念’假说:模型似乎依据属性的概率局部优先级行动,但对真正驱动其决策的属性仅有有限的言语表达能力。

原文摘要 · Abstract (English)

We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.

大模型决策表面信念行为建模可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。