arXiv:2505.17735cs.AI2025-05ACL被引 7

用自动模拟器生成安全数据,让大模型代理更可靠。

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

  • 构建可扩展的威胁模型,精准刻画不安全行为成因。
  • 自动生成大规模安全训练数据,提升真实任务安全表现28.91%。
  • 无需真实危险数据,适合部署在客服、助手等高风险场景。

基于大语言模型(LLM)的智能体在数字助理、自主客服与决策支持系统中广泛应用,其多轮交互与工具调用能力至关重要。然而,用户动态交互、外部工具使用及潜在有害行为带来的安全风险难以保障。为此,我们提出AutoSafe框架,首次实现通过全自动合成数据生成系统性提升智能体安全性。具体而言:1)提出开放可扩展的威胁模型OTS,形式化用户指令、交互上下文与代理行为之间的不安全演化机制;2)构建全自动数据生成流水线,模拟不安全用户行为,通过自我反思生成安全响应,构建大规模、多样且高质量的安全训练数据集,避免采集真实危险数据。我们在合成与真实世界安全基准上进行综合实验,结果表明AutoSafe平均提升安全得分45%,在真实任务上实现28.91%的改进,验证了所学安全策略的泛化能力。该成果推动了面向真实部署的可信赖智能体建设。项目主页已公开:https://auto-safe.github.io/。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents are increasingly deployed in real-world applications such as "digital assistants, autonomous customer service, and decision-support systems", where their ability to "interact in multi-turn, tool-augmented environments" makes them indispensable. However, ensuring the safety of these agents remains a significant challenge due to the diverse and complex risks arising from dynamic user interactions, external tool usage, and the potential for unintended harmful behaviors. To address this critical issue, we propose AutoSafe, the first framework that systematically enhances agent safety through fully automated synthetic data generation. Concretely, 1) we introduce an open and extensible threat model, OTS, which formalizes how unsafe behaviors emerge from the interplay of user instructions, interaction contexts, and agent actions. This enables precise modeling of safety risks across diverse scenarios. 2) we develop a fully automated data generation pipeline that simulates unsafe user behaviors, applies self-reflective reasoning to generate safe responses, and constructs a large-scale, diverse, and high-quality safety training dataset-eliminating the need for hazardous real-world data collection. To evaluate the effectiveness of our framework, we design comprehensive experiments on both synthetic and real-world safety benchmarks. Results demonstrate that AutoSafe boosts safety scores by 45% on average and achieves a 28.91% improvement on real-world tasks, validating the generalization ability of our learned safety strategies. These results highlight the practical advancement and scalability of AutoSafe in building safer LLM-based agents for real-world deployment. We have released the project page at https://auto-safe.github.io/.

大模型安全智能体自动化测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。