首个面向真实场景的智能体安全评估基准,揭示高任务表现与严重安全隐患并存。
BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional Environments
- 构建真实功能环境下的多领域安全风险指令空间,覆盖网页、移动端等四类场景。
- 13个主流智能体平均仅完成不足40%任务且完全合规,任务强表现常伴严重安全违规。
- 融合规则检查与大模型评判的混合评估框架,可量化真实环境中的行为风险。
大型多模态模型(LMMs)使智能体能执行复杂数字与物理任务,但其作为自主决策者部署时存在显著非预期行为安全风险。现有评估受限于低保真环境、模拟API或窄范围任务,缺乏全面性。为此,我们提出BeSafe-Bench(BSB),一个面向功能环境中情境化智能体的行为安全风险评估基准,涵盖网页、移动设备、具身视觉语言模型(VLM)和具身视觉语言代理(VLA)四个代表性领域。通过功能性环境构建多样化指令空间,引入九类安全关键风险,并采用规则检查与大模型作为裁判相结合的混合评估框架,以衡量真实环境影响。对13个主流智能体的评估显示:即使最优者也仅完成少于40%的任务且完全符合安全约束,而强任务表现常伴随严重安全违规。结果凸显了在现实部署前加强安全对齐的紧迫性。
原文摘要 · Abstract (English)
The rapid evolution of Large Multimodal Models (LMMs) has enabled agents to perform complex digital and physical tasks, yet their deployment as autonomous decision-makers introduces substantial unintentional behavioral safety risks. However, the absence of a comprehensive safety benchmark remains a major bottleneck, as existing evaluations rely on low-fidelity environments, simulated APIs, or narrowly scoped tasks. To address this gap, we present BeSafe-Bench (BSB), a benchmark for exposing behavioral safety risks of situated agents in functional environments, covering four representative domains: Web, Mobile, Embodied VLM, and Embodied VLA. Using functional environments, we construct a diverse instruction space by augmenting tasks with nine categories of safety-critical risks, and adopt a hybrid evaluation framework that combines rule-based checks with LLM-as-a-judge reasoning to assess real environmental impacts. Evaluating 13 popular agents reveals a concerning trend: even the best-performing agent completes fewer than 40% of tasks while fully adhering to safety constraints, and strong task performance frequently coincides with severe safety violations. These findings underscore the urgent need for improved safety alignment before deploying agentic systems in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。