arXiv:2411.16251cs.CLcs.LG2024-11被引 1

用可解释的概率编辑法替代黑箱生成器,提升文本分类解释的透明度与稳定性。

Transparent Neighborhood Approximation for Text Classifier Explanation

  • 基于文本上下文的概率编辑生成邻近样本,替代黑箱生成模型。
  • 在两个真实数据集上表现接近生成式方法,且解释更稳定。
  • 全程透明可控,适合需要可解释性的高风险场景应用。

现有研究强调邻域构造在生成模型无关解释中的关键作用,趋势是采用生成模型提升合成样本质量,尤其用于解释文本分类器。这类方法克服了文本非结构化带来的邻域构建挑战,从而提升解释质量。然而,部署的生成器通常基于神经网络,缺乏内在可解释性,引发对解释过程透明度的质疑。为此,本文提出一种基于概率的文本编辑方法,作为黑箱生成器的替代方案。该方法通过基于文本上下文的操纵生成邻近文本。将生成器驱动的邻域构建替换为递归概率编辑后,提出的XPROB(基于概率编辑的解释器)在两个真实数据集上的评估中表现出竞争力。此外,XPROB完全透明且更具可控性的构建过程,相比生成式解释器展现出更优的稳定性。

原文摘要 · Abstract (English)

Recent literature highlights the critical role of neighborhood construction in deriving model-agnostic explanations, with a growing trend toward deploying generative models to improve synthetic instance quality, especially for explaining text classifiers. These approaches overcome the challenges in neighborhood construction posed by the unstructured nature of texts, thereby improving the quality of explanations. However, the deployed generators are usually implemented via neural networks and lack inherent explainability, sparking arguments over the transparency of the explanation process itself. To address this limitation while preserving neighborhood quality, this paper introduces a probability-based editing method as an alternative to black-box text generators. This approach generates neighboring texts by implementing manipulations based on in-text contexts. Substituting the generator-based construction process with recursive probability-based editing, the resultant explanation method, XPROB (explainer with probability-based editing), exhibits competitive performance according to the evaluation conducted on two real-world datasets. Additionally, XPROB's fully transparent and more controllable construction process leads to superior stability compared to the generator-based explainers.

可解释性文本解释概率编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。