提出混合门控注意力机制,提升模型表达能力与训练效率。
Hybrid Gated Attention

- 融合多阶段注意力信息,实现元素级与头间协同门控。
- 在多个基准上优于传统门控注意力,且在不同计算成本下表现最优。
- 引入低秩分解与可学习注意力汇聚点,增强训练稳定性和效率。
门控注意力能有效缓解注意力坍缩并提升表示能力。为进一步拓展其效能-效率权衡边界,我们提出混合门控注意力(HyGA)框架,包含三种门控策略。这些门控利用注意力多阶段的多样化信息,从多个角度协同构建元素级与头级门控,捕捉头内与头间交互。通过混合门控组件,HyGA提供多源调制信号,实现更全面的信息流控制,提升注意力的表征能力。此外,引入低秩矩阵分解和可学习注意力汇聚点,进一步提升训练效率与稳定性。在基于不同骨干网络的广泛基准测试中,实验表明,相比门控注意力,HyGA在训练损失及多种下游任务性能上均有全面提升。在不同计算开销下,其性能均达到最佳,并通过全面的模型分析深化理解。所提方法为更高效、稳定、强大的注意力机制提供了新思路。
原文摘要 · Abstract (English)
Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of gating strategies. Specifically, these gates leverage diverse information from multiple stages of attention, and collaboratively build element-wise/head-wise gating from multiple perspectives, capturing intra-head and cross-head information interactions. Through our hybrid gating components, HyGA could provide multi-source modulation signals, enabling more comprehensive control over information flow and improving the representational capacity of attention. We also introduce low-rank matrix decomposition and learnable attention sink to further enhance training efficiency and stability. In experiments, we evaluate HyGA on widely-used benchmarks based on different backbones. The experimental results show that our HyGA comprehensively improves both training loss and various downstream performances compared with Gated attention. HyGA has also been verified to achieve the best performance at different computation costs, with comprehensive model analyses for better understanding. The proposed HyGA sheds light on a more effective, efficient, and stable attention mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。