arXiv:2506.16402cs.AIcs.CL2025-06AAAI被引 42

首个评估视觉语言模型机器人交互安全性的基准,揭示现有系统缺乏实时风险应对能力。

IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks

  • 构建161个高保真场景,涵盖388种独特安全风险,支持动态交互评估。
  • 实测主流VLM模型在任务中普遍忽视中间步骤风险,安全动作常错序执行。
  • 适合研究具身智能安全、机器人决策可靠性的学者与开发者使用。

由视觉语言模型驱动的具身智能体在日常家务任务中存在规划缺陷,带来显著安全隐患,阻碍其在真实场景中的部署。然而,现有静态、非交互式评估范式无法有效衡量此类交互环境中的动态风险,因无法模拟智能体行为引发的实时风险,且依赖不可靠的事后评估,忽略不安全的中间步骤。为填补这一关键空白,我们提出交互安全性评估:即智能体感知突发风险并按正确流程执行缓解措施的能力。为此,我们构建了IS-Bench,首个面向交互安全的多模态基准,包含161个挑战性场景和388种独特的安全风险,均在高保真模拟器中实现。该基准支持新型过程导向评估,可验证风险缓解动作是否在特定高危步骤前/后执行。对GPT-4o、Gemini-2.5系列等领先VLM的大量实验表明,当前智能体普遍缺乏交互安全意识;尽管引入安全感知思维链可提升表现,但常以牺牲任务完成率为代价。该研究为开发更安全可靠的具身智能系统提供了基础。代码与数据已开源:https://github.com/AI45Lab/IS-Bench。

原文摘要 · Abstract (English)

Flawed planning from VLM-driven embodied agents poses significant safety hazards, hindering their deployment in real-world household tasks. However, existing static, non-interactive evaluation paradigms fail to adequately assess risks within these interactive environments, since they cannot simulate dynamic risks that emerge from an agent's actions and rely on unreliable post-hoc evaluations that ignore unsafe intermediate steps. To bridge this critical gap, we propose evaluating an agent's interactive safety: its ability to perceive emergent risks and execute mitigation steps in the correct procedural order. We thus present IS-Bench, the first multi-modal benchmark designed for interactive safety, featuring 161 challenging scenarios with 388 unique safety risks instantiated in a high-fidelity simulator. Crucially, it facilitates a novel process-oriented evaluation that verifies whether risk mitigation actions are performed before/after specific risk-prone steps. Extensive experiments on leading VLMs, including the GPT-4o and Gemini-2.5 series, reveal that current agents lack interactive safety awareness, and that while safety-aware Chain-of-Thought can improve performance, it often compromises task completion. By highlighting these critical limitations, IS-Bench provides a foundation for developing safer and more reliable embodied AI systems. Code and data are released under https://github.com/AI45Lab/IS-Bench.

具身智能交互安全视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。