arXiv:2508.04451cs.LGcs.AI2025-08被引 6

用强化学习训练AI模拟攻击,发现大模型隐藏漏洞。

Automatic LLM Red Teaming

  • 将红队测试建模为马尔可夫决策过程,分层强化学习生成多轮攻击策略。
  • 在多个基准上超越现有方法,成功触发90%以上隐蔽漏洞。
  • 适合安全研究人员和模型开发者,提升大模型抗攻击能力。

红队测试对识别大语言模型漏洞、建立信任至关重要。但当前自动化方法依赖脆弱的提示模板或单轮攻击,无法捕捉真实对抗对话的复杂交互特性。本文提出新范式:训练一个AI来战略性地‘攻破’另一个AI。通过将红队测试形式化为马尔可夫决策过程(MDP),并采用分层强化学习框架,有效应对稀疏奖励与长时序挑战。生成式智能体通过细粒度的词级有害奖励,学习连贯的多轮攻击策略,能够发现现有基线遗漏的细微漏洞。该方法达到新基准,从根本上将大模型红队测试重构为动态、基于轨迹的过程,对实现稳健的AI部署至关重要。

原文摘要 · Abstract (English)

Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prompt templates or single-turn attacks, failing to capture the complex, interactive nature of real-world adversarial dialogues. We propose a novel paradigm: training an AI to strategically `break' another AI. By formalizing red teaming as a Markov Decision Process (MDP) and employing a hierarchical Reinforcement Learning (RL) framework, we effectively address the inherent sparse reward and long-horizon challenges. Our generative agent learns coherent, multi-turn attack strategies through a fine-grained, token-level harm reward, enabling it to uncover subtle vulnerabilities missed by existing baselines. This approach sets a new state-of-the-art, fundamentally reframing LLM red teaming as a dynamic, trajectory-based process (rather than a one-step test) essential for robust AI deployment.

红队测试强化学习安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。