arXiv:2510.22014cs.CLcs.AI2025-10中稿 · TMLR 2026

揭示大模型对抗后缀迁移性的关键统计规律

Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models

  • 通过分析三类统计特征预测攻击后缀迁移成功率
  • 后缀引发的拒绝方向偏移越大,迁移性越强
  • 适合研究模型安全与对抗攻击的从业者参考

基于离散优化的越狱攻击旨在生成简短、无意义的后缀,附加到输入提示后诱导大语言模型输出违规内容。值得注意的是,这些后缀常具有迁移性——在未经过优化的提示和模型上仍能成功。尽管迁移性现象已被广泛观察,但其发生机制尚缺乏严谨分析。为此,我们在多种实验设置下识别出三个与迁移成功显著相关的统计特性:(1) 无后缀提示激活模型内部拒绝方向的程度;(2) 后缀推动模型远离该方向的强度;(3) 在拒绝方向正交方向上的偏移量大小。相反,提示语义相似性与迁移成功仅弱相关。这些发现提供了对迁移性更精细的理解,并通过干预实验验证了该统计分析可实际提升攻击成功率。

原文摘要 · Abstract (English)

Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed content. Notably, these suffixes are often transferable -- succeeding on prompts and models for which they were never optimized. And yet, despite the fact that transferability is surprising and empirically well-established, the field lacks a rigorous analysis of when and why transfer occurs. To fill this gap, we identify three statistical properties that strongly correlate with transfer success across numerous experimental settings: (1) how much a prompt without a suffix activates a model's internal refusal direction, (2) how strongly a suffix induces a push away from this direction, and (3) how large these shifts are in directions orthogonal to refusal. On the other hand, we find that prompt semantic similarity only weakly correlates with transfer success. These findings lead to a more fine-grained understanding of transferability, which we use in interventional experiments to showcase how our statistical analysis can translate into practical improvements in attack success.

对抗攻击大模型安全迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。