arXiv:2410.17922cs.AI2024-10被引 8

用多智能体框架提升大模型安全防护,兼顾通用性与领域专用性。

Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models

  • 引入多智能体系统,利用外部信息精准判断用户意图并生成安全响应
  • 在通用和化学领域均有效防御越狱攻击,且不降低模型正常功能
  • 适合关注大模型安全与实用平衡的研究者和开发者

随着大语言模型(LLMs)的广泛应用,其安全性日益重要。现有防御方法常面临两大问题:(i) 防护能力不足,尤其在化学等专业领域,因缺乏专业知识导致对恶意查询生成有害回应;(ii) 防御过度,影响模型的通用性和响应能力。为此,我们提出基于多智能体的防御框架G4D,通过利用准确的外部信息,提供对用户意图的无偏分析,并生成基于分析的安全响应指导。在主流越狱攻击和良性数据集上的大量实验表明,G4D能在不损害模型通用功能的前提下,显著增强模型在通用和领域特定场景下对越狱攻击的鲁棒性。

原文摘要 · Abstract (English)

With the extensive deployment of Large Language Models (LLMs), ensuring their safety has become increasingly critical. However, existing defense methods often struggle with two key issues: (i) inadequate defense capabilities, particularly in domain-specific scenarios like chemistry, where a lack of specialized knowledge can lead to the generation of harmful responses to malicious queries. (ii) over-defensiveness, which compromises the general utility and responsiveness of LLMs. To mitigate these issues, we introduce a multi-agents-based defense framework, Guide for Defense (G4D), which leverages accurate external information to provide an unbiased summary of user intentions and analytically grounded safety response guidance. Extensive experiments on popular jailbreak attacks and benign datasets show that our G4D can enhance LLM's robustness against jailbreak attacks on general and domain-specific scenarios without compromising the model's general functionality.

大模型安全多智能体越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。