arXiv:2510.08329cs.CL2025-10

无需种子指令,自动生成多样化攻击提示,提升大模型安全测试效果

AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming

  • 不依赖种子指令,通过角色引导生成自由形式的对抗性提示
  • 在8个主流大模型上实现更高攻击成功率和更强泛化能力
  • 自研验证器高效评估提示危害性,适合模型安全研究人员使用

大语言模型的安全性对可信AI应用至关重要。现有红队测试方法常依赖种子指令,限制了生成对抗性提示的语义多样性。我们提出AutoRed,一种无需种子指令的自由形式对抗提示生成框架,包含两个阶段:(1) 角色引导的对抗指令生成,(2) 反思循环迭代优化低质量提示。为提高效率,引入验证器在不查询目标模型的前提下评估提示危害性。基于AutoRed构建了两个红队测试数据集——AutoRed-Medium和AutoRed-Hard,并评估了八种前沿大模型。结果表明,AutoRed在攻击成功率和泛化性能上均优于现有基线。研究揭示了种子方法的局限性,展示了自由形式红队测试在大模型安全评估中的潜力。相关数据集将于近期开源。

原文摘要 · Abstract (English)

The safety of Large Language Models (LLMs) is crucial for the development of trustworthy AI applications. Existing red teaming methods often rely on seed instructions, which limits the semantic diversity of the synthesized adversarial prompts. We propose AutoRed, a free-form adversarial prompt generation framework that removes the need for seed instructions. AutoRed operates in two stages: (1) persona-guided adversarial instruction generation, and (2) a reflection loop to iteratively refine low-quality prompts. To improve efficiency, we introduce a verifier to assess prompt harmfulness without querying the target models. Using AutoRed, we build two red teaming datasets -- AutoRed-Medium and AutoRed-Hard -- and evaluate eight state-of-the-art LLMs. AutoRed achieves higher attack success rates and better generalization than existing baselines. Our results highlight the limitations of seed-based approaches and demonstrate the potential of free-form red teaming for LLM safety evaluation. We will open source our datasets in the near future.

大模型安全红队测试对抗生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。