从经济学视角揭示GAE的本质机制与缺陷,提出更合理的解释方法CAH。
Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective

- 基于零和博弈经济视角,重构Transformer解释思路。
- 发现GAE仅关注注意力过程,忽略特征层面信息。
- 新方法CAH兼顾过程与特征,适用于含特殊标记的模型。
当前可解释人工智能(XAI)研究过度追求代理指标性能,导致方法本身缺乏合理性和可解释性。本文聚焦广泛影响的GAE方法,揭示其真实工作机制与内在缺陷,提出新的解释范式——累积资产持有(CAH),融合过程与特征双重视角,源于经济零和博弈理论。该方法不仅改进了对注意力机制的理解,还适用于存在特殊标记的模型,突破现有方法局限。研究中采用的模型简化与加法操作分析策略,也为其他XAI研究提供启发。
原文摘要 · Abstract (English)
We observe a phenomenon that current algorithmic research in the field of explainable artificial intelligence primarily pursues better performance on several proxy metrics. On the one hand, these proxy metrics themselves are more or less flawed and cannot properly measure the quality of methods. On the other hand, metric-oriented research approaches often lead to the neglect of the rationality and interpretability of the methods themselves. Explainable artificial intelligence is abbreviated as XAI. The metric-driven research paradigm has resulted in a lack of interpretability of the relevant XAI methods themselves. Accordingly, there is a need for interpretability research on XAI methods, which can be playfully referred to as XXAI. This paper is one of our works on XXAI. This paper takes Generic Attention-model Explainability (GAE), a widely influential model interpretation method , or rather, XAI method that represents an important technical route, as the research object, and explores the real working mechanism and flaws of this method as well as the technical route it represents. Based on the conclusions of this study, it may be necessary to re-examine or verify GAE-related methods and their domain applications. We argue that GAE is an interpretation method that focuses on the attention process. After pointing out the working mechanism and flaws of GAE, we propose Cumulative Asset Holdings (CAH), a more reasonable Transformer interpretation method integrating both process-based and feature-based ideas from an economic zero-sum games perspective. In addition, it is worth noting that our method is applicable to models with special tokens, where existing methods may suffer from limitations. The model simplification research method and the analysis of additive operations adopted in this study may provide inspiration for other research works in XAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。