arXiv:2508.11388cs.CLcs.CV2025-08ACL被引 3

通过优化输入掩码生成神经网络预测解释,无需额外训练模型。

Model Interpretability and Rationale Extraction by Input Mask Optimization

  • 用梯度优化掩码,保留关键输入部分以生成解释。
  • 满足充分性、全面性和紧凑性,提升解释质量。
  • 适用于文本和图像,通用性强,适合研究可解释性者。

随着神经网络在自然语言处理和计算机视觉等领域的快速发展,对黑箱模型预测进行解释的需求日益增长。本文提出一种新方法,通过掩码输入中非指示性部分来生成神经网络预测的抽取式解释。该掩码基于梯度优化,并结合一种新正则化方案,以确保生成解释具备充分性、全面性和紧凑性——这三个特性在自然语言处理中的论据提取领域被广泛认为是理想属性。本方法实现了模型可解释性与论据提取之间的桥梁,证明了无需训练专用模型即可完成论据提取,仅依赖已训练分类器即可实现。此外,该方法同样应用于图像输入,成功生成高质量的图像分类解释,表明自然语言处理中论据提取的条件具有更广泛的适用性。

原文摘要 · Abstract (English)

Concurrent to the rapid progress in the development of neural-network based models in areas like natural language processing and computer vision, the need for creating explanations for the predictions of these black-box models has risen steadily. We propose a new method to generate extractive explanations for predictions made by neural networks, that is based on masking parts of the input which the model does not consider to be indicative of the respective class. The masking is done using gradient-based optimization combined with a new regularization scheme that enforces sufficiency, comprehensiveness and compactness of the generated explanation, three properties that are known to be desirable from the related field of rationale extraction in natural language processing. In this way, we bridge the gap between model interpretability and rationale extraction, thereby proving that the latter of which can be performed without training a specialized model, only on the basis of a trained classifier. We further apply the same method to image inputs and obtain high quality explanations for image classifications, which indicates that the conditions proposed for rationale extraction in natural language processing are more broadly applicable to different input types.

可解释性模型解释输入掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。