arXiv:2606.04778cs.AIcs.CL2026-06

模型生成过程中的任意阶段都可能被短文本干扰,需全程对齐才安全。

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

论文配图:Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories
图 1 · 摘自论文原文
  • 在生成过程中任意步骤注入短文本,都能改变后续输出安全性。
  • 模型内部状态的对齐程度无法预测其抗干扰能力。
  • 通过模拟中途干扰训练模型,可提升对多种攻击的防御力。

安全对齐的大语言模型在推理阶段仍易受干预影响,导致生成有害内容。现有研究将此归因于‘浅层安全’——对齐集中在前几轮输出。我们发现,浅层安全只是更广泛推理时脆弱性的特例:在生成过程任意阶段注入短文本,均可显著改变后续安全行为。此外,模型隐藏层中对拒绝指令的对齐程度,并不能预测其抗干扰能力,表明内部状态不足以决定扰动下的生成表现。为此,我们提出直接在模拟中途扰动的生成轨迹上进行对齐训练,结果表明该方法能有效提升对中段注入攻击的鲁棒性,并泛化至利用早期生成特征的攻击。本工作强调,真正可靠的安全部署必须基于生成过程本身进行训练,而非仅关注输出结果。

原文摘要 · Abstract (English)

Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shallow safety, where alignment concentrates in the first few output tokens. We show that shallow safety is a special case of a broader inference-time vulnerability, in which short token injections at any generation step can substantially alter subsequent safety behavior. We also find that a model's alignment with refusal directions in its hidden states does not predict its robustness to such injection, revealing that internal state alone does not determine generation behavior under perturbation. To address this, we align models directly on generation trajectories constructed by simulating mid-sequence perturbation, and show that this improves robustness to mid-sequence injection and generalizes to attacks that exploit early-token generation. Our work argues that robust safety alignment requires training on the generation process itself, not only its outputs.

大模型安全生成对抗对齐训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。