arXiv:2504.15699cs.AI2025-04IJCAI被引 13

为具身智能体设计安全输入过滤框架,提升防护精度与速度。

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation

  • 构建具身代理专用的安全评测基准与输入过滤架构
  • 检测准确率达94.58%,单次处理仅需0.002秒
  • 适合关注智能体安全的开发者与研究者

具身智能体在众多领域展现出巨大潜力,其行为安全性是实现广泛应用的前提。然而现有研究主要聚焦通用大语言模型的安全性,缺乏针对具身智能体的专门安全评测与输入过滤方法。为此,本文提出一种全新的输入过滤框架,覆盖分类体系定义、数据集构建、模组架构设计、模型训练与严格评估全流程。特别地,我们构建了EAsafetyBench——一个专为具身智能体设计的安全评测基准,用于训练和评估专用模组。此外,提出Pinpoint方案,采用掩码注意力机制实现提示解耦,有效抑制功能提示对安全判断的干扰。在多个基准数据集与模型上的大量实验验证了该方法的有效性:平均检测准确率达94.58%,远超现有最先进方法;单次处理耗时仅0.002秒。

原文摘要 · Abstract (English)

Embodied agents exhibit immense potential across a multitude of domains, making the assurance of their behavioral safety a fundamental prerequisite for their widespread deployment. However, existing research predominantly concentrates on the security of general large language models, lacking specialized methodologies for establishing safety benchmarks and input moderation tailored to embodied agents. To bridge this gap, this paper introduces a novel input moderation framework, meticulously designed to safeguard embodied agents. This framework encompasses the entire pipeline, including taxonomy definition, dataset curation, moderator architecture, model training, and rigorous evaluation. Notably, we introduce EAsafetyBench, a meticulously crafted safety benchmark engineered to facilitate both the training and stringent assessment of moderators specifically designed for embodied agents. Furthermore, we propose Pinpoint, an innovative prompt-decoupled input moderation scheme that harnesses a masked attention mechanism to effectively isolate and mitigate the influence of functional prompts on moderation tasks. Extensive experiments conducted on diverse benchmark datasets and models validate the feasibility and efficacy of the proposed approach. The results demonstrate that our methodologies achieve an impressive average detection accuracy of 94.58%, surpassing the performance of existing state-of-the-art techniques, alongside an exceptional moderation processing time of merely 0.002 seconds per instance.

具身智能安全过滤评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。