arXiv:2607.19914cs.AI2026-07

提出新方法解决长期决策中的风险规划问题,可精确验证结果并支持多种风险偏好。

Long-Term Sequential Decision Making under Risk

  • 用秩-分位数替代目标函数,结合动态规划求解非线性风险目标
  • 在离散回报网格上精确评估策略,提供上下界证书确保结果可信
  • 无需采样或枚举,支持快速参数调整,适合风险敏感决策场景

我们研究基于根式(坚定)风险目标的有限时域马尔可夫决策过程规划,该目标对总回报分布应用一种依赖排名的函数。这类目标在回报分布上是非线性的,通常破坏贝尔曼最优性,因此直接通过情景树枚举进行优化不可行。本文提出ERQDP——一种无需枚举、无需采样的方法,通过精确动态规划求解秩-分位数代理目标;在离散回报网格上利用动态规划精确评估候选策略的回报概率质量函数(PMFs),并给出显式的舍入误差界;通过任意时间循环迭代精炼代理目标,报告目标函数的显式上下界差距(证书),直至离散化预算限制。在多个基准测试中,ERQDP 能返回带证明的解或明确残差间隙,实现风险参数的快速扫描,显著提升运行效率,并支持风险规避与风险追求行为。

原文摘要 · Abstract (English)

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

强化学习风险规划动态规划决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。