为大模型驱动的智能体设计安全评估与对齐框架,提升任务规划安全性。
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
- 构建涵盖2027个日常任务的多灾种安全评测基准
- 引入物理世界安全知识使不安全行为减少8.55%至15.22%
- 适用于需高安全性的具身智能体研发与验证
大型语言模型(LLMs)在提升具身智能体的任务规划能力方面展现出巨大潜力,因其具备强大的推理与理解能力。然而,这些智能体的系统性安全仍是一个未充分探索的领域。本研究提出Safe-BeAl框架,包含安全评估(SafePlan-Bench)与安全对齐(Safe-Align)两部分。SafePlan-Bench建立了一个综合性评测基准,涵盖2,027个日常任务及对应环境,覆盖8类不同危险场景(如火灾危险)。实证分析显示,即使无对抗输入或恶意意图,基于LLM的智能体仍可能表现出不安全行为。为此,我们提出Safe-Align方法,旨在将物理世界安全知识融入基于LLM的具身智能体,同时保持任务特定性能。多场景实验表明,相比基于GPT-4的智能体,Safe-BeAl实现8.55%-15.22%的安全性提升,且保证任务成功完成。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit substantial promise in enhancing task-planning capabilities within embodied agents due to their advanced reasoning and comprehension. However, the systemic safety of these agents remains an underexplored frontier. In this study, we present Safe-BeAl, an integrated framework for the measurement (SafePlan-Bench) and alignment (Safe-Align) of LLM-based embodied agents' behaviors. SafePlan-Bench establishes a comprehensive benchmark for evaluating task-planning safety, encompassing 2,027 daily tasks and corresponding environments distributed across 8 distinct hazard categories (e.g., Fire Hazard). Our empirical analysis reveals that even in the absence of adversarial inputs or malicious intent, LLM-based agents can exhibit unsafe behaviors. To mitigate these hazards, we propose Safe-Align, a method designed to integrate physical-world safety knowledge into LLM-based embodied agents while maintaining task-specific performance. Experiments across a variety of settings demonstrate that Safe-BeAl provides comprehensive safety validation, improving safety by 8.55 - 15.22%, compared to embodied agents based on GPT-4, while ensuring successful task completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。