arXiv:2509.25885cs.AI2025-09被引 14

提出安全风险评估与防御框架,提升具身大模型的物理交互安全性。

SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents

  • 构建四阶段风险模型与三类安全约束,系统识别潜在危险。
  • 在5558样本的SafeMindBench上,主流模型仍频繁出现安全失误。
  • 设计分层安全模块,显著提升安全性且不降低任务完成率。

由大语言模型驱动的具身智能体具备强大的规划能力,但其与物理世界的直接交互带来了安全风险。本文识别出四个可能产生危险的推理阶段:任务理解、环境感知、高层计划生成与底层动作生成,并形式化三类独立的安全约束(事实性、因果性、时间性),以系统刻画潜在的安全违规。基于此风险模型,我们构建了SafeMindBench,一个包含5558个样本的多模态基准,覆盖四大任务类别(Instr-Risk、Env-Risk、Order-Fix、Req-Align),涵盖破坏、伤害、隐私泄露及非法行为等高风险场景。在SafeMindBench上的大量实验表明,当前领先的LLM(如GPT-4o)和广泛使用的具身智能体仍易发生安全关键性失败。为此,我们提出SafeMindAgent,一种集成三个级联安全模块的模块化规划-执行架构,将安全约束融入推理过程。结果表明,SafeMindAgent在强基线基础上显著提升安全率,同时保持相近的任务完成率。SafeMindBench与SafeMindAgent共同提供了严谨的评估体系与实用解决方案,推动具身大模型安全风险的系统性研究与缓解。

原文摘要 · Abstract (English)

Embodied agents powered by large language models (LLMs) inherit advanced planning capabilities; however, their direct interaction with the physical world exposes them to safety vulnerabilities. In this work, we identify four key reasoning stages where hazards may arise: Task Understanding, Environment Perception, High-Level Plan Generation, and Low-Level Action Generation. We further formalize three orthogonal safety constraint types (Factual, Causal, and Temporal) to systematically characterize potential safety violations. Building on this risk model, we present SafeMindBench, a multimodal benchmark with 5,558 samples spanning four task categories (Instr-Risk, Env-Risk, Order-Fix, Req-Align) across high-risk scenarios such as sabotage, harm, privacy, and illegal behavior. Extensive experiments on SafeMindBench reveal that leading LLMs (e.g., GPT-4o) and widely used embodied agents remain susceptible to safety-critical failures. To address this challenge, we introduce SafeMindAgent, a modular Planner-Executor architecture integrated with three cascaded safety modules, which incorporate safety constraints into the reasoning process. Results show that SafeMindAgent significantly improves safety rate over strong baselines while maintaining comparable task completion. Together, SafeMindBench and SafeMindAgent provide both a rigorous evaluation suite and a practical solution that advance the systematic study and mitigation of safety risks in embodied LLM agents.

具身智能安全风险大模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。