arXiv:2507.00026cs.LGcs.AI2025-07

让大模型对抗测试更全面,自动生成多样高危提示。

RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models

  • 用上下文生成+多目标强化学习,自动生成跨领域攻击提示。
  • 在多个评测中提升综合效果,显著增加提示主题多样性。
  • 适合安全研究者和模型开发者用于发现潜在风险漏洞。

随着大语言模型(LLMs)在真实应用中日益作为黑箱组件部署,红队测试已成为识别潜在风险的关键手段。它通过对抗性提示检验LLM,以发现漏洞并提升安全性对齐。理想情况下,有效的红队测试应能适应不断演进的LLM能力,并覆盖广泛的有害主题。然而,现有方法存在两大局限:1)基于主题的方法依赖预收集的有害主题,灵活性与适应性受限;2)无主题方法使用强化学习(RL),但缺乏显式的探索奖励信号,易过度优化单一目标,导致主题多样性下降。为此,我们提出RedTopic,一种新颖的红队测试框架,通过上下文化生成流程、聚合奖励设计和多目标强化学习训练循环,生成主题多样化的对抗性提示。实验表明,RedTopic生成的对抗提示比现有方法更有效且更具多样性,综合评估指标有明显提升。我们认为,RedTopic代表了向更适应性、更主题多样的大模型红队测试迈出的重要一步。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed as black-box components in real-world applications, red teaming has become essential for identifying potential risks. It tests LLMs with adversarial prompts to uncover vulnerabilities and improve safety alignment. Ideally, effective red teaming should be adaptive to evolving LLM capabilities and explore a broad range of harmful topics. However, existing approaches face two limitations: 1) topic-based approaches rely on pre-collected harmful topics, limited in flexibility and adaptivity. 2) topic-free methods use reinforcement learning (RL), but they lack an explicit reward signal for exploration and tend to over-optimize a narrow objective, reducing topic diversity. To address these limitations, we propose RedTopic, a novel red teaming framework that generates topic-diverse adversarial prompts through a contextualized generation pipeline, an aggregate reward design, and a multi-objective RL training loop. Experiments show that RedTopic produces more effective and diverse adversarial prompts than existing methods, with notable improvements in integrated evaluation metrics. We believe RedTopic represents a step toward more adaptive and topic-diverse red teaming for large language models.

红队测试大模型安全对抗生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。