arXiv:2502.05163cs.CLcs.LG2025-02被引 24

用博弈论方法生成高质量多语言安全数据,提升大模型安全性

Enhancing LLM Safety Through a Theoretical Minimax Game Lens

  • 让生成器和分类器互相对抗进化,自动产出生效的安全数据
  • 小模型用生成数据后在英文基准上超主流水平近10%,推理快4.5倍
  • 适合需要多语言安全检测的AI系统研发者快速部署

大语言模型快速发展亟需有效机制以准确区分有害内容与良性内容。尽管英语安全数据集丰富,但其他语言的开源安全数据稀缺,且英语数据集中敏感边缘案例不足,导致模型出现捷径学习和显著误报率。为此,我们提出一种新颖的极小极大强化学习框架,使数据生成器与分类器协同进化,生成高质量合成多语言安全数据。我们从理论上将该互动形式化为极小极大博弈,并严格证明其收敛至纳什均衡。实证评估表明,该合成数据生成方法显著提升分类模型性能,使更小模型在英文基准上超越当前最优水平近10%,推理速度提升4.5倍。该方法为合成数据生成提供了可扩展、高效率的路径,推动更安全、更鲁棒的多语言大模型部署。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) necessitates effective mechanisms to ensure their responsible deployment by accurately distinguishing unsafe content from benign content. While substantial safety datasets are available in English, multilingual safety modeling remains underexplored due to limited open-source safety datasets in other languages. Even within English datasets, safe yet sensitive corner-case content is scarce, leading to shortcut learning by models and non-trivial false-positive rates. To mitigate these issues, we introduce a novel minimax reinforcement learning (RL) framework wherein a data generator and a classifier model co-evolve, facilitating the production of high-quality synthetic multilingual safety data. We theoretically formalize this interaction as a minimax game and rigorously demonstrate convergence to a Nash equilibrium. Empirical evaluations confirm that our synthetic data generation method significantly enhances the classifier model performance, enabling a substantially smaller model to surpass the state-of-the-art by nearly 10% on English benchmarks while achieving 4.5x faster inference speed. These results establish a scalable and efficient methodology for synthetic data generation, advancing the development of safer and more robust multilingual LLM deployments.

LLM安全生成对抗多语言合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。