用影响函数定位关键训练样本和概念,提升NLP模型可解释性。
CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models
- 基于影响函数分析训练样本与概念对预测的影响
- 修复有害样本后性能恢复至基线,无需重新训练
- 适合需要透明决策过程的医疗、金融等高风险领域
近年来,深度学习模型的黑箱特性限制了其在医疗诊断、金融等高风险领域的应用,这些场景中可解释性至关重要。为此,我们提出一种新方法,利用影响函数在样本级和概念级增强NLP模型的可解释性。在CEBaB和Yelp数据集上的实验表明,影响函数能有效识别对模型预测最具影响力的训练样本(包括有益和有害样本)。通过调整这些样本的标签与权重,模型性能可恢复至基线水平,无需重新训练,验证了影响函数在高效数据调试中的价值。此外,概念级分析揭示了概念瓶颈模型(CBM)中显著影响预测的关键概念。修改这些概念会明显改变模型行为,提供清晰的决策过程洞察。
原文摘要 · Abstract (English)
In recent years, the black-box nature of deep learning models has limited their application in high-stakes domains such as medical diagnosis and finance, where interpretability is essential. To address this, we propose a novel approach using influence functions to enhance interpretability in NLP models at both the sample and concept levels. Experiments on CEBaB and Yelp datasets show that influence functions effectively identify the most impactful training samples, both helpful and harmful, on model predictions. By adjusting the labels and weights of these samples, we demonstrate that model performance can be restored to baseline levels without retraining, confirming the value of influence functions for efficient data debugging. Furthermore, our concept-level analysis identifies key concepts within Concept Bottleneck Models (CBM) that significantly affect predictions. Modifying these concepts alters model behavior observably, providing clear insights into the decision process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。