arXiv:2510.06994cs.CRcs.CL2025-10被引 1

用自适应策略生成多轮对抗攻击,全面测试大模型安全性

RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning

  • 设计分层攻击规划器,动态生成多样化对话攻击
  • 多轮攻击使顶尖大模型产生不安全输出,成功率显著
  • 适合大模型安全审计、漏洞挖掘与防御研究者使用

本文提出RedTWIZ:一种自适应且多样化的多轮红队攻击框架,用于评估人工智能辅助软件开发中大型语言模型(LLMs)的鲁棒性。工作基于三大研究方向:(1) 系统化评估大模型对话越狱的鲁棒性;(2) 构建支持组合性、真实性和目标导向的多轮攻击语料库;(3) 设计分层攻击规划器,可自适应地规划、序列化并触发针对特定大模型弱点的攻击。三者融合形成统一框架,实现评估、攻击生成与策略规划一体化,全面揭示大模型的潜在缺陷。大量实验系统评估了整体系统及各组件性能,结果表明多轮对抗攻击可成功诱导前沿大模型生成不安全内容,凸显提升大模型鲁棒性的紧迫性。

原文摘要 · Abstract (English)

This paper presents the vision, scientific contributions, and technical details of RedTWIZ: an adaptive and diverse multi-turn red teaming framework, to audit the robustness of Large Language Models (LLMs) in AI-assisted software development. Our work is driven by three major research streams: (1) robust and systematic assessment of LLM conversational jailbreaks; (2) a diverse generative multi-turn attack suite, supporting compositional, realistic and goal-oriented jailbreak conversational strategies; and (3) a hierarchical attack planner, which adaptively plans, serializes, and triggers attacks tailored to specific LLM's vulnerabilities. Together, these contributions form a unified framework -- combining assessment, attack generation, and strategic planning -- to comprehensively evaluate and expose weaknesses in LLMs' robustness. Extensive evaluation is conducted to systematically assess and analyze the performance of the overall system and each component. Experimental results demonstrate that our multi-turn adversarial attack strategies can successfully lead state-of-the-art LLMs to produce unsafe generations, highlighting the pressing need for more research into enhancing LLM's robustness.

大模型安全红队攻击对抗测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。