揭示了语言模型解释中的词汇与位置偏见,发现不同方法间存在权衡关系。
Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
- 构建无模型依赖的评估框架,系统分析解释方法的词汇和位置偏好
- 在人工与自然数据上均发现词汇与位置偏见呈负相关,高一方低另一方
- 异常解释更易受偏见影响,适合关注解释可信度的研究者阅读
高质量的解释能增强对语言模型和数据的理解。特征归因方法(如集成梯度)作为后处理解释工具,可提供词元级别的洞察。然而,同一输入的不同方法产生的解释差异显著,源于方法内在偏见。用户可能因此质疑其可靠性,而不知情者则可能过度信赖。本文超越表面不一致,通过一种模型与方法无关的框架,使用三项评估指标系统分析两种Transformer模型在人工数据上的伪随机分类任务,以及自然数据上的半控制因果关系检测任务中所表现出的词汇与位置偏见(即解释关注什么、在哪里)。结果发现:词汇偏见与位置偏见之间存在权衡,高得分一方对应低得分另一方;同时,异常解释更易呈现明显偏见。
原文摘要 · Abstract (English)
Good quality explanations strengthen the understanding of language models and data. Feature attribution methods, such as Integrated Gradient, are a type of post-hoc explainer that can provide token-level insights. However, explanations on the same input may vary greatly due to underlying biases of different methods. Users may be aware of this issue and mistrust their utility, while unaware users may trust them inadequately. In this work, we delve beyond the superficial inconsistencies between attribution methods, structuring their biases through a model- and method-agnostic framework of three evaluation metrics. We systematically assess both lexical and position bias (what and where in the input) for two transformers; first, in a controlled, pseudo-random classification task on artificial data; then, in a semi-controlled causal relation detection task on natural data. We find a trade-off between lexical and position biases in our model comparison, with models that score high on one type score low on the other. We also find signs that anomalous explanations are more likely to be biased.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。