为机器人操作安全评估与改进构建了基于规范的基准测试与数据集
MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation

- 基于任务-约束分类法设计200个带接触的家居任务,独立评估安全与成功
- 1000个场景下验证:21%成功操作仍不安全,微调后安全完成率提升至近30%
- 提供8000条带安全标注示范,适合研究机器人安全与强化学习的学者
面向机器人操作的基座模型政策快速发展,但对其是否安全的严谨评估仍不足。我们提出ManiGuard,一个基于规范的安全评估与改进框架,包含ManiGuard-Bench任务套件和配套的安全标注轨迹生成管道。ManiGuard-Bench将六类高接触度家庭任务组织为200个锁定基础任务,依据技能×约束分类法,并独立于任务成功率设定安全规范。每个任务在一种分布内及四种单轴分布外扰动下评估,保持安全规范不变,共形成1000个锁定场景。每次仿真或物理弗兰卡平台上运行的轨迹均由基于LTL$_f$的自动机监控器实时检查,基于物理谓词而非学习分类器或大模型判别。管道结合自动化运动规划与人工遥操作,每步由同一监控器标注,支持安全感知微调;我们公开8000条安全标注示范,每基础任务40条。在超过23,000次滚动中对零样本与微调视觉语言动作模型(VLAs)进行基准测试,发现:(i) 安全必须独立于任务成功评估,6-21%的成功轨迹违反规范;(ii) 在本套件上微调可使安全任务完成率从接近零提升至7.5-29.8%,安全交互行为从16-40%升至51-72%;(iii) 仍存在差距,扩大示范数据无法弥合,21-42%的交互轨迹仍违规,六类任务中有两类在所有策略下安全成功率低于2%,且失败现象在分布偏移和硬件上持续存在。
原文摘要 · Abstract (English)
Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite and a paired safety-annotated trajectory-generation pipeline. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill $\times$ constraint taxonomy, with safety specified independently of task success. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked scenarios. Every rollout is runtime-checked by LTL$_f$-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in simulation and on a physical Franka platform. The pipeline pairs an automated motion-planning generator with human teleoperation, annotated by the same per-step monitor, and directly supports safety-aware fine-tuning; we release 8,000 safety-annotated demonstrations, 40 per base task. Benchmarking zero-shot and fine-tuned VLAs across more than 23,000 rollouts, we find: (i) safety must be evaluated independently of task success, as 6-21% of successful rollouts violate the specification; (ii) fine-tuning on our suite raises safe task completion from near zero to 7.5-29.8% and engaged-and-safe behavior from 16-40% to 51-72%; but (iii) a gap remains that scaling demonstrations does not close, with 21-42% of engaged rollouts still violating, two of six families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。