arXiv:2602.15173cs.AI2026-02ACL

对比大模型在风险决策中的表现,发现推理型更理性,对话型更受表述影响。

Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs

  • 按是否具备数学推理训练,将模型分为理性型与对话型两类
  • 对话型模型对选项描述顺序、正负框架和解释敏感,存在明显描述-历史差距
  • 推理型模型行为接近最优理性,几乎不受表达方式影响,适合高可靠性场景

大型语言模型在决策支持系统或智能体工作流中的应用日益广泛,但其在不确定性下的决策机制仍不清楚。本研究从前景表征(显式表达或结果历史)和决策理由(解释)两个维度,考察了20个前沿开源大模型的风险决策行为,并与人类被试实验及理性期望收益最大化模型进行对比。结果显示,模型可分为两类:推理模型(RMs)和对话模型(CMs)。RMs表现出较高理性,对前景顺序、盈亏框架和解释内容不敏感,无论前景是显式给出还是通过历史结果呈现,行为一致;而CMs显著缺乏理性,略接近人类,对前景顺序、框架和解释敏感,且存在巨大描述-历史差距。开放模型的配对分析表明,区分两者的关键在于是否经过数学推理训练。

原文摘要 · Abstract (English)

The use of large language models either as decision support systems, or in agentic workflows, is rapidly transforming the digital ecosystem. However, the understanding of LLM decision-making under uncertainty remains limited. We study LLM risky choices along two dimensions: (1) prospect representation (based on an explicit representation or outcome history) and (2) decision rationale (explanation). Our study, which involves 20 frontier and open LLMs, is complemented by a matched human subjects experiment, which provides one reference point, while an expected payoff maximizing rational agent model provides another. We find that LLMs cluster into two categories: reasoning models (RMs) and conversational models (CMs). RMs tend towards rational behavior, are insensitive to the order of prospects, gain/loss framing, and explanations, and behave similarly whether prospects are explicit or presented via a history of outcomes. CMs are significantly less rational, slightly more human-like, sensitive to prospect ordering, framing, and explanation, and exhibit a large description-history gap. Paired comparisons of open LLMs suggest that a key factor differentiating RMs and CMs is training for mathematical reasoning.

大模型决策风险偏好推理能力行为差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。