让大模型攻击与防御同步进化,提升安全对齐能力。
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- 用树搜索策略探索多样越狱路径,发现模型漏洞。
- 通过课程强化学习联合训练攻防模型,实现双向优化。
- 适合关注大模型安全与可持续对齐的研究者。
大型语言模型(LLMs)在互联网服务中快速发展,虽带来前所未有的能力,但也放大了社会风险。现有工作多聚焦于孤立的越狱攻击或静态防御,忽视了真实网络环境中威胁与防护的动态互动。为此,我们提出ACE-Safety(Adversarial Co-Evolution for LLM Safety),一种联合优化攻击与防御模型的新框架,无缝集成两项核心创新:(1) 分组感知的策略引导蒙特卡洛树搜索(GS-MCTS),高效探索越狱策略,揭示漏洞并生成多样化对抗样本;(2) 对抗课程树感知的分组策略优化(AC-TGPO),通过课程强化学习,以挑战性样本联合训练攻击与防御LLM,实现稳健的相互提升。在多个基准上的评估表明,该方法优于现有攻防技术,为构建可持续支持负责任人工智能生态系统的模型提供了可行路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either isolated jailbreak attacks or static defenses, neglecting the dynamic interplay between evolving threats and safeguards in real-world web contexts. To mitigate these challenges, we propose ACE-Safety (Adversarial Co-Evolution for LLM Safety), a novel framework that jointly optimize attack and defense models by seamlessly integrating two key innovative procedures: (1) Group-aware Strategy-guided Monte Carlo Tree Search (GS-MCTS), which efficiently explores jailbreak strategies to uncover vulnerabilities and generate diverse adversarial samples; (2) Adversarial Curriculum Tree-aware Group Policy Optimization (AC-TGPO), which jointly trains attack and defense LLMs with challenging samples via curriculum reinforcement learning, enabling robust mutual improvement. Evaluations across multiple benchmarks demonstrate that our method outperforms existing attack and defense approaches, and provides a feasible pathway for developing LLMs that can sustainably support responsible AI ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。