arXiv:2605.12386cs.RO2026-05被引 4

为机器人操作设计可复用的时序安全评估基准,识别任务成功但不安全的行为。

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

论文配图:SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
图 1 · 摘自论文原文
  • 用有限时序逻辑定义安全模板,通过符号化轨迹监测时序安全属性。
  • 在50个家庭任务中测试6个策略,发现强模型仍频繁出现安全违规。
  • 适合关注机器人安全性的研究者,尤其关注时序错误与跨任务泛化。

机器人操作通常以任务完成率评估,但任务成功并不意味着执行安全。许多安全问题具有时序特性:如污染后触碰清洁表面、物体未完全嵌入即释放。我们提出SafeManip,一个基于属性驱动的时序安全评估基准,超越以往仅关注任务完成或状态约束的评估方式。SafeManip利用有限轨迹上的线性时序逻辑(LTLf)定义可复用的安全模板,将观测轨迹映射为符号谓词序列,并通过LTLf监控器进行评估。其属性集涵盖8类操作安全:碰撞与接触安全、抓取稳定性、释放稳定性、交叉污染、动作启动、机制恢复、物体封装与容器访问。模板可针对具体对象、工装、区域或技能实例化,实现跨任务与环境的泛化。我们在6个视觉-语言-动作策略(包括$π_0$、$π_{0.5}$、GR00T及其训练变体)上,对50个RoboCasa365家庭任务进行评估。结果显示,即使表现强劲的模型也常出现不安全行为;任务成功率提升并不必然带来更安全执行:许多成功轨迹仍存在安全违规,且长时序或复杂任务暴露更多违规。SafeManip提供可复用的评估层,用于诊断时序安全失败并衡量超出任务完成的安全成功率。

原文摘要 · Abstract (English)

Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are temporal: a robot may touch a clean surface after contamination or release an object before it is fully inside an enclosure. We introduce SafeManip, a property-driven benchmark to explicitly evaluate temporal safety properties in robotic manipulation, moving beyond prior evaluations that largely focus on task completion or per-state constraint violations. SafeManip defines reusable safety templates over finite executions using Linear Temporal Logic over finite traces (LTLf). It maps observed rollouts to symbolic predicate traces and evaluates them with LTLf-based monitors. Its property suite covers eight manipulation safety categories: collision and contact safety, grasp stability, release stability, cross-contamination, action onset, mechanism recovery, object containment, and enclosure access. Templates can be instantiated with task-specific objects, fixtures, regions, or skills, allowing the same safety specifications to generalize across tasks and environments. We evaluate SafeManip on six vision-language-action policies, including $π_0$, $π_{0.5}$, GR00T, and their training variants, across 50 RoboCasa365 household tasks. Results show that even strong models often behave unsafely. Task-success gains do not reliably translate into safer execution: many successful rollouts remain unsafe, while longer-horizon or more complex tasks expose more violations. SafeManip provides a reusable evaluation layer for diagnosing temporal safety failures and measuring safe success beyond task completion.

机器人安全时序逻辑评估基准操作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。