arXiv:2504.09712cs.CRcs.AI2025-04ACL被引 2

发现大模型安全漏洞在语义等价输入间无法泛化,提出可解释的防御框架

The Structural Safety Generalization Problem

  • 设计语义等价的多轮、多图、翻译攻击,系统测试安全泛化失效
  • 不同输入结构导致安全表现差异显著,暴露关键脆弱性
  • 提出结构重写护栏,提升有害输入拒绝率且不误拒正常输入

大型语言模型的越狱攻击是广泛存在的安全挑战。针对此问题尚未可解,我们聚焦核心失败机制:安全能力在语义等价输入间无法泛化。通过限定攻击需具备可解释性、跨模型迁移性及跨目标迁移性,我们在该框架下开展红队测试,发现了针对多轮对话、多图像输入及翻译攻击的新漏洞。这些攻击在语义上等价于单轮、单图像或未翻译版本,支持系统性对比;结果表明不同输入结构导致不同的安全表现。随后,我们提出结构重写护栏(Structure Rewriting Guardrail),将输入转换为更利于安全评估的结构,显著提升对有害输入的拒绝率,且不过度拒绝良性输入。通过将这一中间挑战(比通用防御更可处理,但对长期安全至关重要)置于研究前沿,我们强调了其作为人工智能安全研究关键里程碑的意义。

原文摘要 · Abstract (English)

LLM jailbreaks are a widespread safety challenge. Given this problem has not yet been tractable, we suggest targeting a key failure mechanism: the failure of safety to generalize across semantically equivalent inputs. We further focus the target by requiring desirable tractability properties of attacks to study: explainability, transferability between models, and transferability between goals. We perform red-teaming within this framework by uncovering new vulnerabilities to multi-turn, multi-image, and translation-based attacks. These attacks are semantically equivalent by our design to their single-turn, single-image, or untranslated counterparts, enabling systematic comparisons; we show that the different structures yield different safety outcomes. We then demonstrate the potential for this framework to enable new defenses by proposing a Structure Rewriting Guardrail, which converts an input to a structure more conducive to safety assessment. This guardrail significantly improves refusal of harmful inputs, without over-refusing benign ones. Thus, by framing this intermediate challenge - more tractable than universal defenses but essential for long-term safety - we highlight a critical milestone for AI safety research.

AI安全越狱攻击结构重写红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。