arXiv:2506.14697cs.CRcs.RO2025-06被引 36

构建多阶段安全评估基准,检测具身智能体对危险指令的响应缺陷。

AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions

  • 设计可扩展的对抗性仿真环境,实现视觉语言模型到具体操作的映射
  • 涵盖45种危险场景、1350个任务和9900条指令,覆盖人、环境、自身风险
  • 从感知到执行分层评估,揭示当前智能体在安全规划上的系统性漏洞

视觉语言模型(VLMs)的融合正推动新一代具身智能体在以人为中心的环境中运行。然而,随着部署扩大,这些系统面临日益增长的安全风险,尤其是在执行危险指令时。现有安全评估基准仍有限:仅覆盖狭窄的危险范围,且主要关注最终结果,忽视了智能体从感知、规划到执行的全过程,从而掩盖了关键故障模式。为此,我们提出SAFE,一个系统评估具身VLM智能体在危险指令下安全性的基准。SAFE包含三部分:SAFE-THOR,一个可扩展的对抗性仿真沙盒,配备通用适配器,将高层VLM输出映射为底层具身控制,支持多种智能体工作流集成;SAFE-VERSE,一个受风险意识启发的任务套件,基于阿西莫夫机器人三定律,包含45个对抗性场景、1350个危险任务和9900条指令,覆盖对人类、环境及智能体自身的风险;SAFE-DIAGNOSE,一种多层级、细粒度的评估协议,衡量智能体在感知、规划和执行各阶段的表现。将SAFE应用于九个先进VLMs和两个具身智能体工作流,我们发现智能体在将危险识别转化为安全规划与执行方面存在系统性失败。研究揭示了当前安全对齐的根本局限,并证明了全面、多阶段评估对于发展更安全具身智能的必要性。

原文摘要 · Abstract (English)

The integration of vision-language models (VLMs) is driving a new generation of embodied agents capable of operating in human-centered environments. However, as deployment expands, these systems face growing safety risks, particularly when executing hazardous instructions. Current safety evaluation benchmarks remain limited: they cover only narrow scopes of hazards and focus primarily on final outcomes, neglecting the agent's full perception-planning-execution process and thereby obscuring critical failure modes. Therefore, we present SAFE, a benchmark for systematically assessing the safety of embodied VLM agents on hazardous instructions. SAFE comprises three components: SAFE-THOR, an extensible adversarial simulation sandbox with a universal adapter that maps high-level VLM outputs to low-level embodied controls, supporting diverse agent workflow integration; SAFE-VERSE, a risk-aware task suite inspired by Asimov's Three Laws of Robotics, comprising 45 adversarial scenarios, 1,350 hazardous tasks, and 9,900 instructions that span risks to humans, environments, and agents; and SAFE-DIAGNOSE, a multi-level and fine-grained evaluation protocol measuring agent performance across perception, planning, and execution. Applying SAFE to nine state-of-the-art VLMs and two embodied agent workflows, we uncover systematic failures in translating hazard recognition into safe planning and execution. Our findings reveal fundamental limitations in current safety alignment and demonstrate the necessity of a comprehensive, multi-stage evaluation for developing safer embodied intelligence.

具身智能安全评估视觉语言模型仿真测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。