arXiv:2507.04117cs.LGcs.CL2025-07

揭示注意力机制隐含的关联偏置,理解其泛化能力来源

Relational inductive biases on attention mechanisms

  • 从几何深度学习视角分析注意力对称性,识别其关系假设
  • 不同注意力层对应不同输入数据关系假设,可分类归纳
  • 适合研究模型泛化与注意力机制设计的学者参考

归纳学习旨在从具体样本中构建通用模型,依赖于影响假设选择并决定泛化能力的先验偏置。本文聚焦于刻画注意力机制中蕴含的关联性归纳偏置,即关于数据元素间潜在关系的假设。基于几何深度学习视角,我们分析了最常见的注意力机制在置换子群下的等变性质,据此提出一种依据其关系偏置的分类方法。在此框架下,我们证明不同注意力层由其对输入数据所假设的底层关系特征所定义。

原文摘要 · Abstract (English)

Inductive learning aims to construct general models from specific examples, guided by biases that influence hypothesis selection and determine generalization capacity. In this work, we focus on characterizing the relational inductive biases present in attention mechanisms, understood as assumptions about the underlying relationships between data elements. From the perspective of geometric deep learning, we analyze the most common attention mechanisms in terms of their equivariance properties with respect to permutation subgroups, which allows us to propose a classification based on their relational biases. Under this perspective, we show that different attention layers are characterized by the underlying relationships they assume on the input data.

注意力机制归纳偏置几何深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。