用监管者视角验证机器人策略的时序安全性,提升自动驾驶行为可靠性。
ROVER: Regulator-Driven Robust Temporal Verification of Black-Box Robot Policies
- 引入监管者闭环机制,通过STL规范评估黑箱策略的时序行为。
- 平均满足率提升43.8%,最差情况违规程度降低,真实机器人路径更平滑。
- 适用于车道保持、加速延迟等复杂时序约束,适合安全关键型机器人验证。
我们提出一种受监管者驱动的黑箱自主机器人策略时序验证方法,灵感来自现实中的认证流程——监管者仅基于可观测行为进行评估,无需访问模型内部。核心是监管者闭环机制,将黑箱策略的执行轨迹与优先级化的信号时序逻辑(STL)安全规范进行对比。这些规范刻画了随时间变化的行为特征,并融入领域知识。我们采用总鲁棒值(TRV)、最大鲁棒值(LRV)和平均违规鲁棒值(AVRV)量化平均性能、最差情况遵守度及平均违规程度,指导针对性重训练与迭代优化。该方法可处理多种时序安全需求(如车道保持、延迟加速、转向平滑性),覆盖虚拟赛车游戏与移动机器人导航两个场景。在两类任务中六项STL规范下,监管引导重训练使满足率平均提升43.8%,半数规范的平均性能(TRV)提高且最坏违规程度(LRV)下降。真实世界验证在TurtleBot3上实现27%的平滑导航满足率提升,路径更顺滑,更符合STL定义的时序安全要求。
原文摘要 · Abstract (English)
We present a novel, regulator-driven approach for the temporal verification of black-box autonomous robot policies, inspired by real-world certification processes where regulators often evaluate observable behavior without access to model internals. Central to our method is a regulator-in-the-loop approach that evaluates execution traces from black-box policies against temporal safety requirements. These requirements, expressed as prioritized Signal Temporal Logic (STL) specifications, characterize behavior changes over time and encode domain knowledge into the verification process. We use Total Robustness Value (TRV) and Largest Robustness Value (LRV) to quantify average performance and worst-case adherence, and introduce Average Violation Robustness Value (AVRV) to measure average specification violation. Together, these metrics guide targeted retraining and iterative model improvement. Our approach accommodates diverse temporal safety requirements (e.g., lane-keeping, delayed acceleration, and turn smoothness), capturing persistence, sequencing, and response across two distinct domains (virtual racing game and mobile robot navigation). Across six STL specifications in both scenarios, regulator-guided retraining increased satisfaction rates by an average of 43.8%, with consistent improvement in average performance (TRV) and reduced violation severity (LRV) in half of the specifications. Finally, real-world validation on a TurtleBot3 robot demonstrates a 27% improvement in smooth-navigation satisfaction, yielding smoother paths and stronger compliance with STL-defined temporal safety requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。