arXiv:2605.14246cs.LGcs.AI2026-05

用近端风险预测提升不完全观测下的安全控制效率

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability

论文配图:Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability
图 1 · 摘自论文原文
  • 构建动作相关的短期风险预测器,替代复杂信念规划
  • 在血糖调控和导航任务中同时优化性能与安全性
  • 适合资源受限、需实时决策的安全关键场景

许多安全关键控制问题可建模为风险敏感的部分可观测马尔可夫决策过程(POMDP),控制器需基于不完整观测做出决策,同时平衡任务表现与安全风险。尽管信念空间规划提供理论保障,但在实际应用中计算成本高且对模型设定敏感。本文提出一种轻量级的风险门控强化学习近似方法,通过构建紧凑的有限历史代理状态,学习动作条件下的近期安全违规预测。该预测风险在两个互补层面使用:一是在价值学习中作为风险惩罚项;二是在决策时作为门控机制,动态插值乐观与保守的集成价值估计。低风险动作更接近奖励导向评估,高风险动作则被更保守地评价。在自动血糖调节和安全约束导航两个领域进行评估,结果表明:在成人与青少年血糖控制队列中,该方法显著改善血糖控制效果并大幅降低运行时间;在Safety-Gym导航基准上,相比无约束RL和多个标准安全强化学习基线,实现了更优的奖励-成本权衡。这些结果表明,在无法进行完整信念空间规划时,动作条件的短期风险可作为有效的局部信号,实现近似的风险敏感型POMDP控制。

原文摘要 · Abstract (English)

Many safety-critical control problems are modeled as risk-sensitive partially observable Markov decision processes, where the controller must make decisions from incomplete observations while balancing task performance against safety risk. Although belief-space planning provides a principled solution, maintaining and planning over beliefs can be computationally costly and sensitive to model specification in practical domains. We propose a lightweight risk-gated reinforcement learning approximation for risk-sensitive control under partial observability. The method constructs a compact finite-history proxy state and learns an action-conditioned predictor of near-term safety violation. This predicted candidate-action risk is used in two complementary ways: as a risk penalty during value learning, and as a decision-time gate that interpolates between optimistic and conservative ensemble value estimates. As a result, low-risk actions are evaluated closer to reward-seeking estimates, while high-risk actions are evaluated more conservatively. We evaluate the approach in two safety-critical partially observable domains: automated glucose regulation and safety-constrained navigation. Across adult and adolescent glucose-control cohorts, the method improves overall glycemic tradeoffs and substantially reduces runtime relative to a belief-space planning baseline. On Safety-Gym navigation benchmarks, it achieves a more favorable reward-cost balance than unconstrained RL and several standard safe-RL baselines. These results suggest that action-conditioned near-term risk can provide an effective local signal for approximate risk-sensitive POMDP control when full belief-space planning is impractical.

安全控制部分可观测强化学习风险感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。