提出可度量可信度的新解释范式,让模型自解释且结果更可信。
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing
- 设计可测量可信度的模型(FMM),使解释可信性可量化优化
- FMM生成的解释接近理论最优可信度,显著优于传统方法
- 适合关注模型透明性与可靠性的研究人员和实践者
随着机器学习在关键场景中的广泛应用,为模型提供可靠解释至关重要。然而,现有解释方法普遍存在可信度不足问题。本文通过构建坚实的可信度度量体系,提出两种新范式:可度量可信度模型(FMM)与自解释机制。FMM通过训练时引入随机掩码等简单修改,使解释的可信度可高效精确测量,从而实现对解释的可信度优化;自解释则探索大模型自我解释的可能性,指出当前模型尚不一致胜任,但提出了可行路径。实验表明,后处理与内在解释方法默认依赖具体模型和任务,而使用FMM后,即使采用相同解释技术,仍能获得稳定可靠的解释。这证明仅通过模型结构的微小调整,即可显著提升解释一致性与可信度,解决了如何提供并确保可信解释的问题。
原文摘要 · Abstract (English)
As machine learning becomes more widespread and is used in more critical applications, it's important to provide explanations for these models, to prevent unintended behavior. Unfortunately, many current interpretability methods struggle with faithfulness. Therefore, this Ph.D. thesis investigates the question "How to provide and ensure faithful explanations for complex general-purpose neural NLP models?" The main thesis is that we should develop new paradigms in interpretability. This is achieved by first developing solid faithfulness metrics and then applying the lessons learned from this investigation to develop new paradigms. The two new paradigms explored are faithfulness measurable models (FMMs) and self-explanations. The idea in self-explanations is to have large language models explain themselves, we identify that current models are not capable of doing this consistently. However, we suggest how this could be achieved. The idea of FMMs is to create models that are designed such that measuring faithfulness is cheap and precise. This makes it possible to optimize an explanation towards maximum faithfulness, which makes FMMs designed to be explained. We find that FMMs yield explanations that are near theoretical optimal in terms of faithfulness. Overall, from all investigations of faithfulness, results show that post-hoc and intrinsic explanations are by default model and task-dependent. However, this was not the case when using FMMs, even with the same post-hoc explanation methods. This shows, that even simple modifications to the model, such as randomly masking the training dataset, as was done in FMMs, can drastically change the situation and result in consistently faithful explanations. This answers the question of how to provide and ensure faithful explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。