arXiv:2608.17386cs.RO2026-08

为机器人操作安全评估与改进构建了基于规范的基准测试与数据集

MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation

论文配图:MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation
图 1 · 摘自论文原文
  • 基于任务-约束分类法设计200个带接触的家居任务,独立评估安全与成功
  • 1000个场景下验证:21%成功操作仍不安全,微调后安全完成率提升至近30%
  • 提供8000条带安全标注示范,适合研究机器人安全与强化学习的学者

面向机器人操作的基座模型政策快速发展,但对其是否安全的严谨评估仍不足。我们提出ManiGuard,一个基于规范的安全评估与改进框架,包含ManiGuard-Bench任务套件和配套的安全标注轨迹生成管道。ManiGuard-Bench将六类高接触度家庭任务组织为200个锁定基础任务,依据技能×约束分类法,并独立于任务成功率设定安全规范。每个任务在一种分布内及四种单轴分布外扰动下评估,保持安全规范不变,共形成1000个锁定场景。每次仿真或物理弗兰卡平台上运行的轨迹均由基于LTL$_f$的自动机监控器实时检查,基于物理谓词而非学习分类器或大模型判别。管道结合自动化运动规划与人工遥操作,每步由同一监控器标注,支持安全感知微调;我们公开8000条安全标注示范,每基础任务40条。在超过23,000次滚动中对零样本与微调视觉语言动作模型(VLAs)进行基准测试,发现:(i) 安全必须独立于任务成功评估,6-21%的成功轨迹违反规范;(ii) 在本套件上微调可使安全任务完成率从接近零提升至7.5-29.8%,安全交互行为从16-40%升至51-72%;(iii) 仍存在差距,扩大示范数据无法弥合,21-42%的交互轨迹仍违规,六类任务中有两类在所有策略下安全成功率低于2%,且失败现象在分布偏移和硬件上持续存在。

原文摘要 · Abstract (English)

Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite and a paired safety-annotated trajectory-generation pipeline. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill $\times$ constraint taxonomy, with safety specified independently of task success. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked scenarios. Every rollout is runtime-checked by LTL$_f$-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in simulation and on a physical Franka platform. The pipeline pairs an automated motion-planning generator with human teleoperation, annotated by the same per-step monitor, and directly supports safety-aware fine-tuning; we release 8,000 safety-annotated demonstrations, 40 per base task. Benchmarking zero-shot and fine-tuned VLAs across more than 23,000 rollouts, we find: (i) safety must be evaluated independently of task success, as 6-21% of successful rollouts violate the specification; (ii) fine-tuning on our suite raises safe task completion from near zero to 7.5-29.8% and engaged-and-safe behavior from 16-40% to 51-72%; but (iii) a gap remains that scaling demonstrations does not close, with 21-42% of engaged rollouts still violating, two of six families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.

机器人安全基准测试强化学习仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。