arXiv:2602.22271cs.LGmath.PR2026-02被引 1

重新定义注意力机制,提出支持令牌提升大模型鲁棒性

Support Tokens, Stability Margins, and a New Foundation for Robust LLMs

  • 将自注意力重构为概率框架,揭示参数的稳定性边界
  • 引入支持令牌概念,训练时加光滑约束项显著增强抗扰动能力
  • 适合关注模型稳定性和鲁棒性优化的研究者

自注意力通常被视为灵活、内容自适应地混合当前词与历史信息的方式。本文在概率框架下重新诠释因果自注意力变压器(即现代基础模型的核心),类似经典PCA扩展为概率PCA。这一重构揭示了变量变换带来的关键结构后果:自注意力参数上出现屏障约束。由此产生的几何结构暴露了注意力映射局部病态的退化边界,可类比支持向量机中的间隔,形成稳定性边界解释。这自然引出‘支持令牌’的概念。进一步证明因果变压器在无限序列上定义了一致的随机过程,为序列建模提供了严格的概率基础。基于此,推导出仅需对标准大模型训练做最小修改的贝叶斯最大后验目标:在交叉熵损失中加入平滑对数屏障惩罚。实验证明,该方法在不牺牲泛化性能的前提下,提升了对输入扰动的鲁棒性,并使学习表征的间隔几何更清晰。

原文摘要 · Abstract (English)

Self-attention is usually described as a flexible, content-adaptive way to mix a token with information from its past. We reinterpret causal self-attention transformers, the backbone of modern foundation models, within a probabilistic framework, much as classical PCA is extended to probabilistic PCA. This reformulation reveals a key structural consequence of the underlying change of variables: a barrier constraint emerges on the parameters of self-attention. The resulting geometry exposes a degeneracy boundary where the attention-induced mapping becomes locally ill-conditioned, yielding a stability-margin interpretation analogous to the margin in support vector machines. This, in turn, naturally gives rise to the concept of support tokens. We further show that causal transformers define a consistent stochastic process over infinite token sequences, providing a rigorous probabilistic foundation for sequence modeling. Building on this view, we derive a Bayesian MAP training objective that requires only a minimal modification to standard LLM training: adding a smooth log-barrier penalty to the usual cross-entropy loss. Empirically, the resulting training objective improves robustness to input perturbations and sharpens the margin geometry of the learned representations without sacrificing out-of-sample accuracy.

大模型鲁棒性注意力机制支持令牌概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。