用对话分析法识别隐蔽攻击,保护大模型安全。
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
- 构建安全对话数据集,训练可插拔的防御模块。
- 多轮攻击成功率降低51.2%,保持模型功能不变。
- 适合关注大模型安全与对抗攻击的研究者。
恶意攻击者可通过多轮对话诱导大型语言模型(LLMs)达成有害目标,带来重大社会安全风险。为此,我们提出一种新型防御机制:多轮对话中的安全推理提取对齐(STREAM)。STREAM在不损害模型功能的前提下,抵御多轮攻击。方法包括构建人类标注的数据集——安全推理多轮对话数据集,并以此微调一个即插即用的安全推理监管模块。该模块能识别隐藏在多轮对话中的恶意意图,并向目标LLM发出潜在风险警告。我们在多个LLM上评估STREAM对主流多轮攻击策略的防御效果。实验结果表明,该方法显著优于现有防御技术,将攻击成功率(ASR)降低51.2%,同时保持了相近的模型能力。
原文摘要 · Abstract (English)
Malicious attackers can exploit large language models (LLMs) by engaging them in multi-turn dialogues to achieve harmful objectives, posing significant safety risks to society. To address this challenge, we propose a novel defense mechanism: SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues (STREAM). STREAM defends LLMs against multi-turn attacks while preserving their functional capabilities. Our approach involves constructing a human-annotated dataset, the Safety Reasoning Multi-turn Dialogues dataset, which is used to fine-tune a plug-and-play safety reasoning moderator. This model is designed to identify malicious intent hidden within multi-turn conversations and alert the target LLM of potential risks. We evaluate STREAM across multiple LLMs against prevalent multi-turn attack strategies. Experimental results demonstrate that our method significantly outperforms existing defense techniques, reducing the Attack Success Rate (ASR) by 51.2%, all while maintaining comparable LLM capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。