多智能体辩论提升机器人任务规划安全性,避免误拒安全指令。
MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning
- 用多个大模型智能体辩论指令安全性,由评判器打分并投票。
- 在虚拟家庭环境中实现90%以上危险任务拒收率,安全任务误拒率低。
- 无需训练,适合需高安全性的家庭服务机器人部署。
确保具身AI代理在任务规划中的安全性对现实世界部署至关重要,尤其在家庭环境中存在危险指令时风险显著。现有方法或因偏好对齐训练导致计算成本高,或因单一安全提示造成过度拒绝。为此,我们提出MADRA,一种无需训练的多智能体辩论风险评估框架,通过集体推理增强安全意识而不牺牲任务表现。MADRA利用多个基于大模型的智能体辩论给定指令的安全性,由关键评判器根据逻辑严谨性、风险识别、证据质量和清晰度评分。通过迭代讨论与共识投票,显著降低误拒率,同时保持对危险任务的高敏感性。此外,我们引入层次化认知协同规划框架,整合安全、记忆、规划与自我演化机制,通过持续学习提升任务成功率。我们还构建了SafeAware-VH数据集,用于虚拟家庭环境中的安全感知任务规划,包含800条标注指令。在AI2-THOR和VirtualHome上的大量实验表明,该方法在超过90%的危险任务上实现拒收,且安全任务误拒率极低,优于现有方法在安全性和执行效率上的表现。本工作提供了一种可扩展、模型无关的可信具身智能解决方案。
原文摘要 · Abstract (English)
Ensuring the safety of embodied AI agents during task planning is critical for real-world deployment, especially in household environments where dangerous instructions pose significant risks. Existing methods often suffer from either high computational costs due to preference alignment training or over-rejection when using single-agent safety prompts. To address these limitations, we propose MADRA, a training-free Multi-Agent Debate Risk Assessment framework that leverages collective reasoning to enhance safety awareness without sacrificing task performance. MADRA employs multiple LLM-based agents to debate the safety of a given instruction, guided by a critical evaluator that scores responses based on logical soundness, risk identification, evidence quality, and clarity. Through iterative deliberation and consensus voting, MADRA significantly reduces false rejections while maintaining high sensitivity to dangerous tasks. Additionally, we introduce a hierarchical cognitive collaborative planning framework that integrates safety, memory, planning, and self-evolution mechanisms to improve task success rates through continuous learning. We also contribute SafeAware-VH, a benchmark dataset for safety-aware task planning in VirtualHome, containing 800 annotated instructions. Extensive experiments on AI2-THOR and VirtualHome demonstrate that our approach achieves over 90% rejection of unsafe tasks while ensuring that safe-task rejection is low, outperforming existing methods in both safety and execution efficiency. Our work provides a scalable, model-agnostic solution for building trustworthy embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。