动态调控策略,让学习到的智能体既安全又高效达成目标
Follow the STARs: Dynamic $ω$-Regular Shielding of Learned Policies
- 用可调参数的STARs框架动态控制干预强度
- 在移动机器人上验证了对概率策略的正确性保障能力
- 适合需要实时适应变化规范的物理系统应用
本文提出一种新型动态后屏蔽框架,用于在预先计算的概率策略上强制执行全类ω-正则正确性属性。这标志着从传统安全屏蔽(确保永不发生坏事)向同时保障活锁性(确保最终发生好事)的范式转变。核心是基于策略模板的自适应运行时屏蔽器(STARs),利用宽容型策略模板实现最小干扰的后屏蔽。其主要特性是动态控制干扰机制,通过可调参数在形式化要求与任务行为间实时权衡:必要时可加强干预,平时则允许优化策略选择。此外,STARs支持运行时适应规范变更或执行器故障,特别适用于网络物理系统。我们在移动机器人基准测试中验证了该方法在增量更新ω-正则属性下对学习策略的可控干扰能力。
原文摘要 · Abstract (English)
This paper presents a novel dynamic post-shielding framework that enforces the full class of $ω$-regular correctness properties over pre-computed probabilistic policies. This constitutes a paradigm shift from the predominant setting of safety-shielding -- i.e., ensuring that nothing bad ever happens -- to a shielding process that additionally enforces liveness -- i.e., ensures that something good eventually happens. At the core, our method uses Strategy-Template-based Adaptive Runtime Shields (STARs), which leverage permissive strategy templates to enable post-shielding with minimal interference. As its main feature, STARs introduce a mechanism to dynamically control interference, allowing a tunable enforcement parameter to balance formal obligations and task-specific behavior at runtime. This allows to trigger more aggressive enforcement when needed, while allowing for optimized policy choices otherwise. In addition, STARs support runtime adaptation to changing specifications or actuator failures, making them especially suited for cyber-physical applications. We evaluate STARs on a mobile robot benchmark to demonstrate their controllable interference when enforcing (incrementally updated) $ω$-regular correctness properties over learned probabilistic policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。