提出新框架诊断大模型在物理任务中的隐蔽安全风险。
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
- 构建双测试框架,分别评估拒绝危险指令与生成安全计划的能力。
- 13个主流大模型在情境风险识别上失败率超60%,暴露隐蔽风险盲区。
- 适合关注机器人决策安全、智能体可靠性的研究者和开发者。
大型语言模型(LLMs)正被广泛用于具身智能体的决策,但现有安全评估多依赖粗粒度成功率和特定领域设置,难以诊断模型失效原因。为此,本文提出SAFEL框架,系统评估LLMs在具身决策中的物理安全性。该框架包含两项核心能力测试:(1)通过命令拒绝测试评估模型拒绝对抗危险指令的能力;(2)通过计划安全测试评估生成安全可执行计划的能力,并将后者分解为目标理解、状态转移建模、动作排序等模块,实现故障的细粒度诊断。为支撑该框架,我们构建了EMBODYGUARD基准,基于PDDL的942个由LLM生成的场景,涵盖明显恶意与情境性危险指令。对13个先进LLMs的评估显示,尽管多数模型能拒绝显性危险指令,但在预判和缓解隐蔽情境风险方面表现不佳。结果揭示了当前大模型在具身安全推理中的关键缺陷,为更精准的模块化改进提供了基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail. This obscures our understanding of embodied safety and limits the selective deployment of LLMs in high-risk physical environments. We introduce SAFEL, the framework for systematically evaluating the physical safety of LLMs in embodied decision making. SAFEL assesses two key competencies: (1) rejecting unsafe commands via the Command Refusal Test, and (2) generating safe and executable plans via the Plan Safety Test. Critically, the latter is decomposed into functional modules, goal interpretation, transition modeling, action sequencing, enabling fine-grained diagnosis of safety failures. To support this framework, we introduce EMBODYGUARD, a PDDL-grounded benchmark containing 942 LLM-generated scenarios covering both overtly malicious and contextually hazardous instructions. Evaluation across 13 state-of-the-art LLMs reveals that while models often reject clearly unsafe commands, they struggle to anticipate and mitigate subtle, situational risks. Our results highlight critical limitations in current LLMs and provide a foundation for more targeted, modular improvements in safe embodied reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。