arXiv:2604.17803cs.AIcs.LG2026-04

用对抗竞赛生成高质量对话数据,提升大模型安全编码能力。

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

论文配图:Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition
图 1 · 摘自论文原文
  • 让攻防双方团队互动竞争,自动产生多样复杂对话
  • 生成1.97万条多轮对话,使代码生成准确率提升超29%
  • 适合研究大模型安全对齐与低成本数据构建的学者

后训练大语言模型需要多样且高质量的数据,但在低资源领域和多轮对话场景中,这类数据稀有且昂贵。现有方法如众包或合成生成常导致数据质量或多样性不足。我们提出对抗竞技场(Adversarial Arena),将数据生成设为对抗任务:攻击方设计提示,防御方生成回应。多团队间交互竞争自然催生出多样且复杂的对话数据。我们组织了来自美国和欧洲顶尖高校的10支学术团队参与该竞赛,聚焦网络安全领域的大模型安全对齐问题,共生成19,683条多轮对话。在该数据集上微调开源模型后,在CyberSecEval-Instruct上安全代码生成准确率提升18.47%,在CyberSecEval-MITRE上提升29.42%。

原文摘要 · Abstract (English)

Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.

对抗生成对话数据安全对齐众包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。