arXiv:2608.04732eess.SYcs.AI2026-08

将不确定性估计、安全过滤与经验回放一体化,提升机器人导航安全性与鲁棒性。

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

论文配图:Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
图 1 · 摘自论文原文
  • 用不确定性动态调整障碍物形状,实时优化安全边界。
  • 执行动作比名义动作更能降低平均代价(7.63±0.44)和信念误差(3.52±0.55 cm)。
  • 适合关注安全强化学习、高可靠性控制的科研与工程人员。

安全演员-评论家控制常将屏障滤波、不确定性估计和经验回放视为独立模块,尽管它们均影响学习与控制所用数据。本文提出一种集成架构:不确定性估计更新控制屏障函数中的障碍物几何,滤波干预与估计残差决定经验回放优先级,评论家从实际执行动作而非名义动作中学习。在二维机器人导航任务中,面对被污染的障碍物测量,对比六种组件匹配配置,在相同训练预算、随机种子、传感器流、探索策略和扰动下进行评估。测试包含中等后训练测试、十一级感知噪声扫描及乘子达6.0的极端压力测试。集成配置在极端测试中五次评估均无接触且成功抵达目标,平均代价为7.63±0.44,障碍物信念均方根误差为3.52±0.55厘米。不确定性估计消融实验也无接触,但仅四次成功,平均代价8.96±2.08,信念误差11.08±1.23厘米。有限训练界限明确回放暴露程度,稳健屏障条件说明所需估计误差与可行性假设。结果支持在该基准上耦合估计、安全过滤与回放;更广泛的安全性与收敛性结论需进一步研究。

原文摘要 · Abstract (English)

Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier function, filter interventions and estimation residuals determine replay priority, and the critic learns from the executed rather than nominal action. We instantiate the architecture on a two-dimensional robot-navigation task with corrupted obstacle measurements and compare six component-matched configurations under common training budgets, random seeds, sensor streams, exploration, and disturbances. Evaluation includes a moderate post-training test, an eleven-level perception-noise sweep, and an exploratory extreme-stress test at multiplier $6.0$. In the extreme test, the integrated configuration recorded no contacts and reached the goal in all five evaluation seeds. Its mean cost was $7.63\pm0.44$ and its obstacle-belief root-mean-square error was $3.52\pm0.55$ cm. The uncertainty-estimation ablation also recorded no contacts but reached the goal in four of five seeds, with mean cost $8.96\pm2.08$ and belief error $11.08\pm1.23$ cm. A finite-training bound clarifies replay exposure, and a robust barrier condition states the required estimation-error and feasibility assumptions. The results support coupling estimation, safety filtering, and replay on this benchmark; broader safety and convergence claims require further study.

安全强化学习不确定性估计经验回放机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。