arXiv:2602.23073math.STcs.AI2026-02

提出加速风险敏感的强化学习决策方法,理论保证高效安全。

Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds

  • 用辅助变量构建CVaR新边界,实现风险价值函数快速估算
  • 在简化信念马尔可夫决策过程上计算上下界,保持原问题一致性
  • 支持动作剪枝,大幅提速且确保安全策略不被误删

在部分可观测环境中进行风险敏感决策是人工智能中的核心挑战,对构建可靠自主智能体至关重要。该问题的正式框架是部分可观测马尔可夫决策过程(POMDP),其中通过风险度量作用于价值函数引入风险敏感性,条件风险价值(CVaR)是一个关键指标。然而,求解一般POMDP在计算上是不可行的,近似方法依赖于昂贵的未来轨迹模拟。本文提出一个理论框架,用于在具有严格性能保证的前提下加速CVaR价值函数评估。我们基于辅助随机变量Y与目标变量X的分布和密度函数关系,推导出新的CVaR边界,这些边界产生可解释的浓度不等式,并在分布差异趋零时收敛。在此基础上,我们建立可从简化信念-MDP中计算的CVaR价值函数上下界,兼容对转移动态的一般简化。我们在粒子信念-MDP框架内设计了带有概率保证的边界估计器,并利用它们实现动作消除:在简化模型下表明次优的动作可安全剔除,同时保持与原始POMDP的一致性。多组实验验证,该边界能有效区分安全与危险策略,在简化模型下实现显著计算加速。

原文摘要 · Abstract (English)

Risk-averse decision-making under uncertainty in partially observable domains is a central challenge in artificial intelligence and is essential for developing reliable autonomous agents. The formal framework for such problems is the partially observable Markov decision process (POMDP), where risk sensitivity is introduced through a risk measure applied to the value function, with Conditional Value-at-Risk (CVaR) being a particularly significant criterion. However, solving POMDPs is computationally intractable in general, and approximate methods rely on computationally expensive simulations of future agent trajectories. This work introduces a theoretical framework for accelerating CVaR value function evaluation in POMDPs with formal performance guarantees. We derive new bounds on the CVaR of a random variable X using an auxiliary random variable Y, under assumptions relating their cumulative distribution and density functions; these bounds yield interpretable concentration inequalities and converge as the distributional discrepancy vanishes. Building on this, we establish upper and lower bounds on the CVaR value function computable from a simplified belief-MDP, accommodating general simplifications of the transition dynamics. We develop estimators for these bounds within a particle-belief MDP framework with probabilistic guarantees, and employ them for acceleration via action elimination: actions whose bounds indicate suboptimality under the simplified model are safely discarded while ensuring consistency with the original POMDP. Empirical evaluation across multiple POMDP domains confirms that the bounds reliably separate safe from dangerous policies while achieving substantial computational speedups under the simplified model.

强化学习风险敏感POMDPCVaR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。