arXiv:2410.18541cs.CLcs.AI2024-10被引 4

重新证明注意力权重能解释模型输出,提出可计算的有效注意力机制。

On Explaining with Attention Matrices

  • 提出有效注意力机制,分离出对预测真正关键的注意力成分。
  • 实验证明有效注意力在4个NLP数据集上具有因果解释力。
  • 适用于需要上下文理解的NLP任务,帮助理解模型决策过程。

本文探讨了变换器模型中注意力权重(AW)与预测结果之间可能存在的解释性关联。尽管早期研究认为注意力权重具有解释价值,但近期研究提出形式论证和实证证据表明其不具解释性。本文指出这些形式论证存在错误,并提出一种高效注意力计算方法,可有效分离出在需上下文信息的任务中起解释作用的注意力矩阵成分。实验表明,有效注意力在多种指标下具备因果解释力(提供最小必要且充分条件),且其矩阵为可计算的概率分布。研究在四个数据集上验证了该方法的多个性质,支持其在解释注意力模型行为中的重要性。

原文摘要 · Abstract (English)

This paper explores the much discussed, possible explanatory link between attention weights (AW) in transformer models and predicted output. Contrary to intuition and early research on attention, more recent prior research has provided formal arguments and empirical evidence that AW are not explanatorily relevant. We show that the formal arguments are incorrect. We introduce and effectively compute efficient attention, which isolates the effective components of attention matrices in tasks and models in which AW play an explanatory role. We show that efficient attention has a causal role (provides minimally necessary and sufficient conditions) for predicting model output in NLP tasks requiring contextual information, and we show, contrary to [7], that efficient attention matrices are probability distributions and are effectively calculable. Thus, they should play an important part in the explanation of attention based model behavior. We offer empirical experiments in support of our method illustrating various properties of efficient attention with various metrics on four datasets.

注意力机制模型解释NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。