通过对抗训练提升大模型在敏感领域的价值一致性
Adversarial Alignment: Ensuring Value Consistency in Large Language Models for Sensitive Domains
- 设计攻防框架:攻击者生成争议问题,执行者输出一致回应,评判者过滤质量
- 在中英文双语数据集上优于主流模型,尤其在政治社会议题上表现更稳
- 适合需要安全可控的场景,如政策咨询、公共对话系统
随着大语言模型广泛应用,其在种族、社会与政治等敏感领域中的偏见和价值不一致问题日益突出。本文提出一种对抗对齐框架,通过持续预训练、指令微调和对抗训练提升模型在敏感领域中的价值一致性。对抗训练中,攻击者生成争议性问题,执行者生成价值一致的回答,评判者负责筛选并保障响应质量。此外,我们构建了中英文双语评估数据集,并训练出一个价值一致的大语言模型(VC-LLM)。实验结果表明,VC-LLM在中英文测试中均优于现有主流模型,验证了方法的有效性。注意:本文包含部分具有冒犯性或有害内容的模型示例。
原文摘要 · Abstract (English)
With the wide application of large language models (LLMs), the problems of bias and value inconsistency in sensitive domains have gradually emerged, especially in terms of race, society and politics. In this paper, we propose an adversarial alignment framework, which enhances the value consistency of the model in sensitive domains through continued pre-training, instruction fine-tuning and adversarial training. In adversarial training, we use the Attacker to generate controversial queries, the Actor to generate responses with value consistency, and the Critic to filter and ensure response quality. Furthermore, we train a Value-Consistent Large Language Model, VC-LLM, for sensitive domains, and construct a bilingual evaluation dataset in Chinese and English. The experimental results show that VC-LLM performs better than the existing mainstream models in both Chinese and English tests, verifying the effectiveness of the method. Warning: This paper contains examples of LLMs that are offensive or harmful in nature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。