arXiv:2410.13785cs.CLcs.AI2024-10被引 3

通过多样化对比模式提升大模型对齐效果,增强抗攻击能力

PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment

  • 从提示、模型、流程三方面引入六种无需额外标注的对比策略
  • 实验表明该方法显著优于现有方法,对齐更全面
  • 适合关注模型安全与泛化能力的研究者

大语言模型对齐通常通过偏好对比输出对训练模型以符合人类偏好。传统方法如RLHF和RLAIF依赖有限的对比模式,如模型变体或解码温度差异。这种单一性导致对齐不全面,且模型易受越狱攻击。为此,本文研究如何构建更全面、多样化的对比模式以提升偏好数据质量(RQ1),并验证对比模式多样化对模型对齐的影响(RQ2)。针对RQ1,提出PopAlign框架,在提示、模型、流水线三个层面集成多样化对比模式,引入六种无需额外反馈标注的对比策略。针对RQ2,通过系统实验验证,PopAlign显著优于现有方法,实现更全面的对齐。

原文摘要 · Abstract (English)

Alignment of large language models (LLMs) involves training models on preference-contrastive output pairs to adjust their responses according to human preferences. To obtain such contrastive pairs, traditional methods like RLHF and RLAIF rely on limited contrasting patterns, such as varying model variants or decoding temperatures. This singularity leads to two issues: (1) alignment is not comprehensive; and thereby (2) models are susceptible to jailbreaking attacks. To address these issues, we investigate how to construct more comprehensive and diversified contrasting patterns to enhance preference data (RQ1) and verify the impact of the diversification of contrasting patterns on model alignment (RQ2). For RQ1, we propose PopAlign, a framework that integrates diversified contrasting patterns across the prompt, model, and pipeline levels, introducing six contrasting strategies that do not require additional feedback labeling procedures. Regarding RQ2, we conduct thorough experiments demonstrating that PopAlign significantly outperforms existing methods, leading to more comprehensive alignment.

大模型对齐对比学习安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。