arXiv:2507.08270cs.AIcs.CR2025-07被引 8

用强化学习统一提升语言模型代理的安全性,防用户和工具双重威胁。

Agent Safety Alignment via Reinforcement Learning

  • 设计沙箱环境与三类标签体系,识别并应对用户与工具的威胁。
  • 在多个基准上验证,安全防护能力显著提升且不影响正常任务性能。
  • 适合关注智能体安全部署的研究者与开发者参考。

具备工具调用能力的自主大语言模型代理带来了超越传统对话滥用的新安全风险。这些代理因可执行外部操作,易受用户发起的攻击(如恶意提示)和工具自身引发的威胁(如被攻陷工具的恶意输出)。本文提出首个面向工具使用型代理的统一安全对齐框架,通过结构化推理与沙箱强化学习,同时应对两类威胁。我们引入包含良性、恶意、敏感三类的双通道分类体系,并构建策略驱动决策模型。框架采用自定义沙箱环境,模拟真实工具执行过程,支持细粒度奖励设计。在Agent SafetyBench、InjecAgent和BFCL等公开及自建基准上的广泛评估表明,安全对齐后的代理在抵御安全威胁方面表现优异,同时保持良好任务效用。结果证明,安全性与有效性可协同优化,为自主大语言模型代理的可信部署奠定基础。

原文摘要 · Abstract (English)

The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents, empowered to execute external functions, are vulnerable to both user-initiated threats (e.g., adversarial prompts) and tool-initiated threats (e.g., malicious outputs from compromised tools). In this paper, we propose the first unified safety-alignment framework for tool-using agents, enabling models to handle both channels of threat via structured reasoning and sandboxed reinforcement learning. We introduce a tri-modal taxonomy, including benign, malicious, and sensitive for both user prompts and tool responses, and define a policy-driven decision model. Our framework employs a custom-designed sandbox environment that simulates real-world tool execution and allows fine-grained reward shaping. Through extensive evaluations on public and self-built benchmarks, including Agent SafetyBench, InjecAgent, and BFCL, we demonstrate that our safety-aligned agents significantly improve resistance to security threats while preserving strong utility on benign tasks. Our results show that safety and effectiveness can be jointly optimized, laying the groundwork for trustworthy deployment of autonomous LLM agents.

智能体安全强化学习大模型工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。