用三维度数据提升大模型安全,有效降低恶意攻击成功率。
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
- 构建三维度风险评估框架,覆盖词汇、恶意意图和越狱策略
- 生成2.6万条有害指令与合规回复,使攻击成功率下降20%
- 适合大模型安全研究者及对齐训练工程师使用
大型语言模型在自然语言处理任务中表现优异,但仍易生成有害内容或被用于恶意目的。尽管已有安全对齐数据集通过监督微调缓解此类风险,但现有数据集往往缺乏全面的风险覆盖,主要关注词汇多样性而忽视其他关键维度。为此,我们提出一种系统性分析框架,从词汇多样性、恶意意图和越狱策略三个核心维度衡量对齐数据集的风险覆盖度。我们进一步设计了TRIDENT自动化流水线,利用基于角色的零样本大模型生成技术,产出跨三维度的多样化指令。每条有害指令均配有伦理对齐响应,形成两个数据集:包含26,311个样本的TRIDENT-Core,以及18,773个样本的TRIDENT-Edge。在Llama 3.1-8B上使用TRIDENT-Edge进行微调,相比在WildBreak数据集上微调的最佳基线模型,平均危害分降低14.29%,攻击成功率下降20%。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploited for malicious purposes. Although safety alignment datasets have been introduced to mitigate such risks through supervised fine-tuning (SFT), these datasets often lack comprehensive risk coverage. Most existing datasets focus primarily on lexical diversity while neglecting other critical dimensions. To address this limitation, we propose a novel analysis framework to systematically measure the risk coverage of alignment datasets across three essential dimensions: Lexical Diversity, Malicious Intent, and Jailbreak Tactics. We further introduce TRIDENT, an automated pipeline that leverages persona-based, zero-shot LLM generation to produce diverse and comprehensive instructions spanning these dimensions. Each harmful instruction is paired with an ethically aligned response, resulting in two datasets: TRIDENT-Core, comprising 26,311 examples, and TRIDENT-Edge, with 18,773 examples. Fine-tuning Llama 3.1-8B on TRIDENT-Edge demonstrates substantial improvements, achieving an average 14.29% reduction in Harm Score, and a 20% decrease in Attack Success Rate compared to the best-performing baseline model fine-tuned on the WildBreak dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。