用生成式方法让文档分类模型的决策变得可解释。
DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
- 基于扩散模型生成贴近真实数据的反事实图像。
- 在3个数据集和3种模型上验证了生成图像的有效性与真实性。
- 适合关注AI决策透明性的研究人员和开发者。
随着黑箱AI决策系统在现代文档处理流程中日益普及,提升其透明度与可靠性变得至关重要,尤其在高风险场景中,模型中的偏见或虚假关联可能导致严重后果。文档图像分类是此类流程中的关键环节,尽管广泛应用,却难以解释。现有方法多依赖特征重要性图,但解释性差,无法揭示模型学习的全局特征。本文提出DocVCE,一种基于扩散模型的生成式反事实解释方法,结合分类器引导生成符合分布的视觉反事实图像,并通过分层块级细化寻找最接近目标图像的优化结果。我们在三个文档分类数据集(RVL-CDIP、Tobacco3482、DocLayNet)和三种模型(ResNet、ConvNeXt、DiT)上进行了严格评估,采用有效性、接近度和真实性等标准验证了方法的有效性。据作者所知,这是首个探索文档图像分析中生成式反事实解释的工作。
原文摘要 · Abstract (English)
As black-box AI-driven decision-making systems become increasingly widespread in modern document processing workflows, improving their transparency and reliability has become critical, especially in high-stakes applications where biases or spurious correlations in decision-making could lead to serious consequences. One vital component often found in such document processing workflows is document image classification, which, despite its widespread use, remains difficult to explain. While some recent works have attempted to explain the decisions of document image classification models through feature-importance maps, these maps are often difficult to interpret and fail to provide insights into the global features learned by the model. In this paper, we aim to bridge this research gap by introducing generative document counterfactuals that provide meaningful insights into the model's decision-making through actionable explanations. In particular, we propose DocVCE, a novel approach that leverages latent diffusion models in combination with classifier guidance to first generate plausible in-distribution visual counterfactual explanations, and then performs hierarchical patch-wise refinement to search for a refined counterfactual that is closest to the target factual image. We demonstrate the effectiveness of our approach through a rigorous qualitative and quantitative assessment on 3 different document classification datasets -- RVL-CDIP, Tobacco3482, and DocLayNet -- and 3 different models -- ResNet, ConvNeXt, and DiT -- using well-established evaluation criteria such as validity, closeness, and realism. To the best of the authors' knowledge, this is the first work to explore generative counterfactual explanations in document image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。