通过压力测试发现人形机器人安全滤波器在复杂环境下的失效模式。
Adversarial Stress Testing of SPARK Humanoid Safety Filters

- 复现并测试6种安全滤波器在不同干扰下的表现
- 部分方法减少碰撞步数,部分更精准追踪目标
- 适合评估人形机器人安全性的研究者和开发者
人形机器人部署困难,因其高维身体、多重碰撞约束,且需在人员和障碍物附近运行。安全滤波器可在名义控制动作可能违反避障约束时进行修正。然而,名义基准得分无法全面反映滤波器在更难环境中的行为。本文通过在MuJoCo中复现SPARK基准案例G1SportMode_D1_WG_SO_v1,对RSSA、RSSS、SSA、CBF、PFM和SMA六种方法在受控随机种子下进行评估,并构建后处理流水线,将原始日志转换为目标追踪、最小距离和碰撞步数指标。结果显示,某些方法更精确追踪目标,另一些则更有效减少碰撞步数。压力测试进一步表明,在障碍物密集、距离估计噪声和延迟障碍信息条件下,安全行为会发生变化。这些发现提示,人形机器人自主性应超越名义性能评估,采用能暴露故障模式的指标以确保部署前安全性。
原文摘要 · Abstract (English)
Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and obstacles. Safety filters help by modifying a nominal control action when it may violate collision-avoidance constraints. Still, nominal benchmark scores do not fully show how these filters behave in harder environments. In this work, we study the robustness of SPARK humanoid safety filters through replication and stress testing. We replicate the SPARK benchmark case G1SportMode_D1_WG_SO_v1 in MuJoCo and evaluate RSSA, RSSS, SSA, CBF, PFM, and SMA under controlled random seeds. We also built a post-processing pipeline that converts raw SPARK logs into goal-tracking, minimum-distance, and collision-step metrics. Our results show that some methods track the goal more closely, while others reduce collision steps more effectively. The stress tests further indicate that safety behavior can change under obstacle crowding, noisy distance estimates, and delayed obstacle information. These findings suggest that humanoid autonomy should be evaluated beyond nominal performance, using metrics that expose failure modes before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。