提出可平滑切换安全与性能行为的自适应策略架构,兼顾安全性与学习连续性。
Safe Online Learning via Smooth Safety-Structured Policy Composition

- 将安全监控与干预嵌入动作生成过程,实现动态响应
- 在多个连续控制基准上实现强安全约束且学习过程平滑
- 已在真实小车-摆杆系统验证,适合实际部署场景
安全在线强化学习要求策略在遵守安全约束的同时保持平滑的优化动态。现有方法通常依赖严格的安全干预,导致系统交互与学习中出现不连续;或采用软约束形式,虽保持学习平滑但安全保证有限。本文提出AutoSafe,一种具备安全感知能力的策略架构,将结构化安全监控与干预直接集成到动作生成过程中。该设计支持性能驱动与安全保护行为之间的平滑、风险相关过渡,从而实现连续的在线交互与学习动态。在一系列连续控制基准测试中,实验结果表明其在不牺牲学习平滑性的前提下实现了强安全约束。进一步在真实物理小车-摆杆系统上验证了AutoSafe的有效性,证明其在现实世界中实现安全在线学习的可行性。
原文摘要 · Abstract (English)
Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics. Existing approaches typically rely on either strict safety enforcement via action interventions, which introduce discontinuities in system interaction and learning, or soft safety constraint formulations, which preserve smooth learning but provide limited safety assurance. We propose AutoSafe, a safety-aware policy architecture that integrates structured safety monitoring and intervention directly into the action generation process. This design enables smooth, risk-dependent transitions between performance-driven and safety-preserving behaviors, resulting in continuous online interaction and learning dynamics. Empirical results across a suite of continuous-control benchmarks demonstrate strong safety enforcement without sacrificing learning smoothness. We further validate AutoSafe on a physical cart-pole system, highlighting its practical effectiveness for safe online learning in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。