arXiv:2608.17574cs.AIcs.SY2026-08中稿 · IJCAI

让智能体根据不确定性动态调整谨慎程度,保障决策安全

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

  • 用后验分布控制鲁棒性范围,随学习逐步放松保守策略
  • 安全价值介于最坏情况与完全已知之间,越学越接近最优
  • 适合对安全性要求高的大模型系统实时决策

智能体在环境未知时应多谨慎?我们提出RATTL(风险对抗总收益学习),将谨慎程度与认知不确定性绑定:智能体持有对未知动态的贝叶斯后验,并针对一个以该后验为参数的Wasserstein模糊集进行规划,模糊半径是后验的单调函数。随着证据积累,半径收缩,行为连续地在最坏情况鲁棒性与无风险收益最大化之间过渡。设计基于熵值风险度量的对偶性,将风险水平选择转化为模糊半径选择。我们证明在暂态和紧致条件下规划问题有解,并建立安全夹逼定理:RATTL值介于无知鲁棒值与全知最优值之间,且当后验集中时间隙消失。在典型二元危险实例中,诱导准则等价于由后验熵设定的条件风险价值。一个示例显示智能体延迟执行高效动作,直至达到明确识别阈值。RATTL旨在保障运行时安全,适用于包括大模型在内的不确定环境下行动的智能体。

原文摘要 · Abstract (English)

How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.

强化学习安全决策不确定性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。