arXiv:2604.18976cs.CL2026-04ACL被引 2

用多智能体网络自动生成越狱提示,更高效发现大模型漏洞

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

论文配图:STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
图 1 · 摘自论文原文
  • 构建策略-响应双重网络,将攻击搜索空间结构化
  • 攻击成功率更高,计算成本更低,优于现有方法
  • 可解释性强,适合安全测试与模型防御研究者

尽管大型语言模型(LLMs)应用广泛,但仍易受越狱提示攻击,导致产生有害或不当回应。本文提出STAR-Teaming,一种新型黑盒自动化红队框架,能有效生成此类攻击提示。该方法结合多智能体系统(MAS)与策略-响应双重网络,通过网络驱动优化采样有效攻击策略。该基于网络的方法将难以处理的高维嵌入空间重构为可管理结构,带来两大优势:提升对LLM战略漏洞的可解释性,并通过语义社区组织搜索空间,避免重复探索。实验表明,STAR-Teaming显著优于现有方法,在更低计算成本下实现更高攻击成功率(ASR)。大量实验验证了双重网络的有效性与可解释性。代码已开源:https://github.com/selectstar-ai/STAR-Teaming-paper。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. This paper introduces STAR-Teaming, a novel black-box framework for automated red teaming that effectively generates such prompts. STAR-Teaming integrates a Multi-Agent System (MAS) with a Strategy-Response Multiplex Network and employs network-driven optimization to sample effective attack strategies. This network-based approach recasts the intractable high-dimensional embedding space into a tractable structure, yielding two key advantages: it enhances the interpretability of the LLM's strategic vulnerabilities, and it streamlines the search for effective strategies by organizing the search space into semantic communities, thereby preventing redundant exploration. Empirical results demonstrate that STAR-Teaming significantly surpasses existing methods, achieving a higher attack success rate (ASR) at a lower computational cost. Extensive experiments validate the effectiveness and explainability of the Multiplex Network. The code is available at https://github.com/selectstar-ai/STAR-Teaming-paper.

红队测试越狱攻击多智能体可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。