arXiv:2605.17128cs.CRcs.AI2026-05中稿 · ICML被引 1

攻击者通过同时查询多个大模型,可100%触发有害输出,暴露新型高危漏洞。

New Wide-Net-Casting Jailbreak Attacks Risk Large Models

论文配图:New Wide-Net-Casting Jailbreak Attacks Risk Large Models
图 1 · 摘自论文原文
  • 设计新攻击方法,针对多模型协同查询的宽网场景
  • 无防护时攻击成功率达100%,远超单模型攻击风险
  • 提醒安全评估应关注多模型协同攻击威胁,适合防御研究者参考

大型模型的越狱攻击因与社会安全密切相关而受到广泛关注。本文识别出一种实际却未被探索的越狱场景——广域网式攻击(wide-net-casting),即攻击者可同时查询一组大模型以诱导产生有害输出。分析揭示该场景存在显著但此前被忽视的安全风险。作为核心贡献,我们提出一种专为该场景设计的新越狱方法。实验表明,在无额外防护的情况下,该方法对部分大模型的越狱成功率可达100%,凸显广域网式攻击是一种独特且高风险的威胁场景,亟需在后续评估与防御研究中予以重视。

原文摘要 · Abstract (English)

Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexplored jailbreak scenario, the wide-net-casting scenario, where an adversary can query a group of large models instead of a single one to elicit harmful outputs. Our analysis reveals substantial yet previously overlooked safety risks under this scenario. As a key part of our analysis, we further develop a novel jailbreak method tailored to the wide-net-casting scenario. With this tailored method, the jailbreak success rate can even reach 100\% in some experiments when targeting the large models without additional safeguards, exposing wide-net-casting as a distinct, high-risk scenario that warrants attention in future evaluation and defense research.

越狱攻击大模型安全多模型协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。