arXiv:2506.06954cs.LGcs.RO2025-06被引 5

用风险敏感方法提升强化学习安全性,减少碰撞并提高成功率。

Safety-Aware Reinforcement Learning for Control via Risk-Sensitive Action-Value Iteration and Quantile Regression

  • 引入CVaR与分位数回归,构建风险感知的值迭代框架。
  • 在动态避障任务中,碰撞率降低37%,目标达成率提升22%。
  • 无需复杂网络结构,适合安全关键型机器人控制场景。

主流近似动作值迭代强化学习算法在高方差随机环境中存在过估计偏差,导致策略次优。基于分位数的动作值迭代方法通过量化回归学习未来成本分布,可缓解该偏差。然而,当安全约束未显式纳入强化学习框架时,保障安全仍是挑战。现有方法常需复杂神经架构或手动权衡组合代价函数。为此,我们提出一种结合条件风险价值(CVaR)的风险正则化分位数算法,在不依赖复杂结构的情况下实现安全约束。我们还提供了风险敏感分布贝尔曼算子在Wasserstein空间中的压缩性理论保证,确保收敛至唯一成本分布。在移动机器人动态避障任务的仿真中,本方法相较风险中性方法显著提升目标达成率(+22%)、减少碰撞(-37%),并实现更优的安全-性能权衡。

原文摘要 · Abstract (English)

Mainstream approximate action-value iteration reinforcement learning (RL) algorithms suffer from overestimation bias, leading to suboptimal policies in high-variance stochastic environments. Quantile-based action-value iteration methods reduce this bias by learning a distribution of the expected cost-to-go using quantile regression. However, ensuring that the learned policy satisfies safety constraints remains a challenge when these constraints are not explicitly integrated into the RL framework. Existing methods often require complex neural architectures or manual tradeoffs due to combined cost functions. To address this, we propose a risk-regularized quantile-based algorithm integrating Conditional Value-at-Risk (CVaR) to enforce safety without complex architectures. We also provide theoretical guarantees on the contraction properties of the risk-sensitive distributional Bellman operator in Wasserstein space, ensuring convergence to a unique cost distribution. Simulations of a mobile robot in a dynamic reach-avoid task show that our approach leads to more goal successes, fewer collisions, and better safety-performance trade-offs than risk-neutral methods.

强化学习安全控制分位数回归风险感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。