arXiv:2606.00970cs.AIcs.LG2026-06

即使无风险偏好,最优控制也会自然产生前景理论特征。

Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States

  • 在灾难状态附近,最优策略自动倾向保守或冒险,取决于趋势方向。
  • 推导出损失厌恶系数的闭式解,与数值结果相关性达0.999。
  • 适用于多种噪声类型,且模型无关,适合研究理性决策机制的人看。

我们研究带有吸收型灾难状态的马尔可夫决策过程中的风险中性控制问题。尽管奖励为线性、无效用弯曲、无概率加权或框架依赖,标准贝尔曼最优性仍表现出三种前景理论特征:价值函数呈S形(灾难附近凸,远场凹)、内生损失敏感度系数λ*(S) > 1,以及反射效应政策反转。在495种配置下,正漂移(增长)时最优策略在灾难附近选择安全动作,即使其即时期望收益更低;负漂移(衰退)时则选择冒险动作,尽管其即时期望损失更高。我们推导出渐近损失厌恶平台值ȳ的闭式表达式,仅依赖胜率p、收益不对称比r = |Δℓ/Δw|和折扣因子β,与数值解拟合优度达R² = 0.999。该机制无需收益不对称。在三组不对称水平下,当r=1.25时,不对称对ȳ>1的贡献中位数为4.6%,升至r=2时达13.9%,且边界贡献始终超过不对称贡献。这些现象在表格Q-learning中依然成立(增长与衰退阶段相关系数分别为0.98和1.00),在高斯、重尾t₃及偏斜正态噪声(最大达步长50%)下也保持稳定,安全通道噪声下预测误差<0.41%,双通道或风险通道噪声下<9.6%。结果表明,吸收型失败状态是生成前景理论特征的充分结构性机制。

原文摘要 · Abstract (English)

We study risk-neutral control in Markov decision processes with an absorbing catastrophic state. Even though rewards are linear and the agent has no utility curvature, probability weighting, or framing dependence, standard Bellman optimality produces three prospect-theory-like signatures: an S-shaped value-function profile (convex near catastrophe, concave in the far field), an endogenous loss-sensitivity coefficient $λ^*(S) > 1$, and a reflection-effect policy reversal. Across 495 configurations, the optimal policy plays safe near catastrophe in positive-drift (growth) regimes despite the risky action's higher immediate expected value, and plays risky near catastrophe in negative-drift (decline) regimes despite the safe action's lower immediate expected loss. We derive a closed-form expression for the asymptotic loss-aversion plateau $\barλ$ that depends only on win probability $p$, payoff asymmetry $r = |Δ_\ell/Δ_w|$, and discount factor $β$, and matches numerical solutions to $R^2 = 0.999$. The mechanism does not require asymmetric payoffs. Across a sweep of $(p,β)$ at three asymmetry levels, the asymmetry share of $\barλ$ above unity has median 4.6% at $r = 1.25$ and rises to 13.9% at $r = 2$, with the boundary contribution exceeding the asymmetry contribution in every cell tested. The phenomena persist under tabular Q-learning (a model-free agent reproduces $V^*$ at correlation 0.98 in growth and 1.00 in decline) and under stochastic transitions with Gaussian, heavy-tailed Student-$t_3$, and asymmetric skew-normal noise up to 50% of the step size, where the asymptotic plateau tracks the closed-form prediction within 0.41% for safe-channel noise and within 9.6% for risky-channel or both-channel noise. These results identify absorbing failure states as a sufficient structural mechanism for prospect-theory-like behavior under optimal control.

强化学习决策理论风险偏好马尔可夫决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。