用对抗竞赛生成高质量对话数据,提升大模型安全编码能力。
Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

- 让攻防双方团队互动竞争,自动产生多样复杂对话
- 生成1.97万条多轮对话,使代码生成准确率提升超29%
- 适合研究大模型安全对齐与低成本数据构建的学者
后训练大语言模型需要多样且高质量的数据,但在低资源领域和多轮对话场景中,这类数据稀有且昂贵。现有方法如众包或合成生成常导致数据质量或多样性不足。我们提出对抗竞技场(Adversarial Arena),将数据生成设为对抗任务:攻击方设计提示,防御方生成回应。多团队间交互竞争自然催生出多样且复杂的对话数据。我们组织了来自美国和欧洲顶尖高校的10支学术团队参与该竞赛,聚焦网络安全领域的大模型安全对齐问题,共生成19,683条多轮对话。在该数据集上微调开源模型后,在CyberSecEval-Instruct上安全代码生成准确率提升18.47%,在CyberSecEval-MITRE上提升29.42%。
原文摘要 · Abstract (English)
Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。