arXiv:2607.11920econ.EMcs.AI2026-07

用主观期望效用敏感度衡量大模型决策一致性,揭示其在不确定情境下的理性程度。

Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making

论文配图:Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
图 1 · 摘自论文原文
  • 引入敏感度参数α,通过软最大选择模型量化对主观期望效用的遵循程度。
  • 实证发现大模型在保险与赌博任务中呈现结构化α差异,但信念与效用参数难精确估计。
  • 方法适用于评估大模型在小样本、高不确定性场景下的决策合理性,适合认知建模研究者。

在标签结果稀缺、昂贵或受运气干扰的情况下,评估不确定性下的决策极具挑战。本文将主观期望效用(SEU)最大化视为基准标准,提出一种分级度量——SEU敏感度,用于衡量代理对这一标准的符合程度。采用带有敏感度参数α的软最大选择模型,对α及信念和效用参数(β, δ)的可识别性进行理论分析,并在Stan中通过先验预测检验、参数恢复和基于模拟的校准(SBC)验证。在仅含不确定选择的模型m₀中,给定期望效用向量η,α可被清晰识别并精确恢复,而(β, δ)仅弱信息:后验分布几乎不收缩且集中在β-δ权衡上。在扩展模型m₁中,δ理论上可通过无β的风险块实现可识别,但在实际样本量下估计改进微乎其微(匹配计数时置信区间宽度减少不足1%),该块也未带来α精度提升。这两个现象表明:可识别性不等于实际精确估计;可识别性对有限样本精度无说明力。即使联合后验信息薄弱,边际SBC仍可通过——我们对此边界进行了精确界定。两组应用(GPT-4o与Claude 3.5 Sonnet,分别处理保险理赔分诊与埃尔斯伯格型瓮问题,以采样温度为调节杠杆)在真实大模型选择数据上端到端运行,检测到四个子情形中两个存在结构性α差异。

原文摘要 · Abstract (English)

Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of an agent's conformity to it. The vehicle is a softmax choice model with a sensitivity parameter $α$ on SEU-valued alternatives; the contribution is a sequence of identifiability results for $α$ and for belief and utility parameters $(β, δ)$, validated in Stan via prior predictive checks, parameter recovery, and simulation-based calibration (SBC), with finite-sample caveats intact. In the uncertain-choice-only model $m_0$, $α$ is identifiable given the expected-utility vector $η$ and sharply recovered, while $(β, δ)$ are only weakly informed: the posterior barely contracts and concentrates on a $β$-$δ$ trade-off. In the extended model $m_1$, $δ$ becomes identifiable in principle via a $β$-free risky block, but its practical recovery gain at realistic sample sizes is negligible (matched-count CI-width reduction under 1%), and that block yields no detected $α$-precision gain at matched choice count. These are two distinct phenomena: for $δ$, identifiability does not imply precise estimability at realistic $n$; for $α$, identifiability is silent about what governs finite-$n$ precision. Marginal SBC passes for both models even where the joint posterior is weakly informed -- a demarcation we make precise. A two-by-two application (GPT-4o and Claude 3.5 Sonnet, each on insurance-claims triage and Ellsberg-style urns, with sampling temperature as the lever) runs end-to-end on real LLM choice data, detecting a structured comparative $α$ effect in two of four cells.

大模型决策行为建模贝叶斯推断风险偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。