arXiv:2509.07324cs.CLcs.AI2025-09EMNLP被引 3

通过一步信念传播提升小模型注意力分布,缓解注意力坍缩问题

Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation

  • 用一步信念传播注入多跳依赖关系,优化自注意力机制
  • 实验显示深层网络熵崩溃减轻,全局依赖度维持在适配任务水平
  • 对小规模模型效果显著,适合资源受限场景使用

基于Transformer的自注意力机制是现代语言模型的核心,但常出现注意力坍缩现象,即关注集中在少数词元上,无法捕捉长程依赖。为此,我们提出自注意力一步信念传播(SAOBP)精炼框架,通过信念传播过程注入多跳关系。为解释和量化这些交互,引入全局词元依赖(GTD),用于捕获注意力图中多跳连接的相对贡献。实证结果表明,SAOBP可有效防止深层网络中的熵坍缩,并自适应地将GTD维持在任务适宜水平,从而提升模型性能。尤为重要的是,在小规模模型中观察到显著收益,凸显其在资源受限场景下提升推理质量的潜力。

原文摘要 · Abstract (English)

Transformer-based self-attention mechanism serves as the core of modern language models, yet it often suffers from localization, where attentions collapse onto a limited subset of tokens and fail to capture long-range dependencies. To address this issue, we propose Self-Attention One-step Belief Propagation (SAOBP), a refinement framework that injects multi-hop relationships through a belief propagation process. To interpret and quantify these interactions, we introduce Global Token Dependency (GTD) that captures the relative contribution of multihop connections within the attention graph. Empirical results indicate that SAOBP helps prevent entropy collapse in deeper layers and adaptively maintains GTD at task-appropriate levels, thereby supporting improvements in model performance. Importantly, we observe competitive gains in small-scale models, highlighting its potential for improving inference quality in resource-constrained scenarios.

自注意力小模型信念传播注意力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。