arXiv:2606.04978cs.CLcs.CY2026-06

LLM在风险决策中看似像人,实则机制不同。

Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game

论文配图:Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game
图 1 · 摘自论文原文
  • 用圣彼得堡游戏测试模型决策机制
  • 多数模型虽出价有限,但行为逻辑不符人类
  • 需关注机制一致性而非仅看结果相似

大型语言模型在风险决策任务中看似谨慎,但表面的谨慎并不等于与人类决策机制对齐。本文以圣彼得堡游戏为可控测试平台,该经典悖论中期望收益无限,但人类通常愿支付有限金额。我们评估了28个LLM,采用结构化提示套件,包括原游戏、扰动截断、重复博弈、资金数额和职业身份的变体,以及要求模型以人类视角推理的提示,还对比了基础模型与指令微调版本。在原游戏中,多数模型生成有限出价,呈现类似人类的风险行为;但控制变量实验显示,模型行为往往转为条件性且计算理性的策略,偏离人类模式。人类引导提示和指令微调虽降低出价并减少部分异常,但多数机制层面响应模式仍无显著变化。结果表明,风险决策中的行为对齐可能是表层的:模型可产出人类似的行为,却不具备人类一致的决策机制。因此,高风险场景下对LLM决策的评估应超越结果相似性,深入考察机制一致性。

原文摘要 · Abstract (English)

LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.

风险决策机制对齐大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。