arXiv:2607.14543cs.ROcs.AI2026-07被引 1

测试视觉语言模型在复杂动作中的空间安全风险,发现任务完成但违规频发。

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

论文配图:SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents
图 1 · 摘自论文原文
  • 构建507个带空间关系的可执行评估样本,区分空间与非空间场景
  • 7个模型任务成功率超90%但过程安全合规率不足60%
  • 适合研究机器人安全、具身智能与多模态推理的开发者

视觉语言模型(VLMs)正被广泛用于具身智能体的决策核心,使机器人能够理解视觉场景、遵循语言指令并规划多步动作。然而,在家庭环境中,安全不仅依赖于物体识别,更取决于动作如何随时间改变物理场景。现有具身安全评估主要关注静态风险识别、拒绝危险指令或最终任务完成,对由支撑、包含、邻近等空间关系引发的过程级安全失效研究不足。为此,我们提出SAFE RELBENCH,一个面向空间关系的具身安全基准,包含507个可执行评估样本,其中248个涉及空间关系,259个为非空间对照样本。用该基准评估7个开源与闭源VLM驱动的具身智能体,发现任务成功与过程级安全合规间存在显著差距:模型常完成任务却违反过程安全约束。与以往基准不同,SAFE RELBENCH显式测试智能体在高风险动作前是否满足安全条件,将空间关系作为具身安全评估的核心维度。结果表明,安全具身智能不仅需要更强感知与规划能力,还需可靠推理对象关系如何影响交互过程中的风险。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions. In household environments, however, safety depends not only on recognizing objects, but also on how actions change the physical scene over time. Existing embodied safety evaluations largely focus on static risk recognition, unsafe instruction refusal, or final-state task completion. As a result, process-level safety failures induced by spatial relations such as support, containment, and proximity remain insufficiently studied. To address this gap, we introduce SAFERELBENCH, a spatial-relation-aware safety benchmark with 507 executable evaluation samples, including 248 spatial-relation samples and 259 non-spatial control samples. Using SAFERELBENCH to evaluate seven open- and closed-source VLM-driven embodied agents, we find a substantial gap between task success and process-level safety compliance: models often complete the requested task while violating process-level safety constraints. Unlike prior benchmarks, SAFERELBENCH explicitly tests whether agents satisfy safety conditions before risk-prone actions, making spatial relations a core dimension in embodied safety assessment. More broadly, our results show that safe embodied intelligence requires not only stronger perception and planning, but also reliable reasoning about how object relations shape risk during interaction.

具身智能安全评估空间关系视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。