用预测安全过滤器让机器人走路更稳,不撞墙也不摔倒。
Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering

- 后置式安全过滤:实时预测碰撞并优化脚位,不改训练策略。
- 仿真与实机验证:复杂环境中事故率大幅下降,动作偏差小。
- 适合需高安全性的机器人控制场景,如救灾或医疗巡检。
强化学习策略可实现动态足式运动,但缺乏避免训练中未覆盖安全约束的能力。大规模离线安全学习难以覆盖所有边缘情况。现有安全框架要么依赖简化模型无法分析全身行为,要么需要保守的恢复控制器而降低任务性能。我们提出一种预测性安全过滤器,对强化学习策略输入的接触点进行后处理。当预测到碰撞时,基于全物理模型的采样优化器异步搜索更安全的接触序列,同时利用学习的价值函数估算长程回报。三个算法组件(采样接触的几何投影、动量增强更新、复制交换)使优化在非连续接触环境下可行。我们在四足机器人上验证了该方法,在密集杂乱环境中实现了仿真与真实世界的显著安全提升,且对原始输入的偏离极小。
原文摘要 · Abstract (English)
Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent during training. Large-scale offline safe learning is impractical for covering all edge cases. Existing safety frameworks either rely on reduced-order models that cannot reason about whole-body behaviors or require conservative recovery controllers that degrade task performance. We propose a predictive safety filter that post-hoc filters the nominal contact locations fed to the RL policy. When a collision is predicted, a sampling-based optimizer asynchronously searches for safer contact sequences using a full-physics model, while a learned value function bootstraps long-horizon returns. Our three algorithmic components (geometric projection of sampled contacts, momentum-augmented updates, and replica-exchange) make the optimization tractable in a discontinuous contact landscape. We validate the filter on a quadruped robot in dense, cluttered environments, both in simulation and in the real world, showing substantial reductions in safety violations with minimal deviation from the nominal input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。