arXiv:2510.00683cs.CV2025-10

用分割模型引导原型学习,让解释更准确可靠。

ProtoMask: Segmentation-Guided Prototype Learning

  • 以图像分割结果裁剪输入,限定原型注意力区域。
  • 在三个细粒度分类数据集上均达领先性能。
  • 适合需要可信赖解释的医疗、工业检测场景。

近年来,可解释人工智能(XAI)受到广泛关注。基于原型案例推理的方法在提升可解释性方面展现出潜力,但通常依赖额外的后处理显著性技术来解释学习到的原型语义,而这些技术的可靠性和质量常受质疑。为此,我们研究利用主流图像分割基础模型,提升嵌入空间与输入空间映射的真实性。通过将显著性图计算区域限制在预定义的语义图像块内,降低可视化不确定性。为感知整幅图像信息,我们使用每个生成的分割掩码的边界框对图像进行裁剪,每张掩码对应一个独立输入,构成名为ProtoMask的新模型架构。我们在三个主流细粒度分类数据集上进行了广泛实验,采用多种评估指标,全面分析了模型的可解释性特征。与多个主流模型对比表明,本模型在性能上具有竞争力,并展现出独特的可解释性优势。

原文摘要 · Abstract (English)

XAI gained considerable importance in recent years. Methods based on prototypical case-based reasoning have shown a promising improvement in explainability. However, these methods typically rely on additional post-hoc saliency techniques to explain the semantics of learned prototypes. Multiple critiques have been raised about the reliability and quality of such techniques. For this reason, we study the use of prominent image segmentation foundation models to improve the truthfulness of the mapping between embedding and input space. We aim to restrict the computation area of the saliency map to a predefined semantic image patch to reduce the uncertainty of such visualizations. To perceive the information of an entire image, we use the bounding box from each generated segmentation mask to crop the image. Each mask results in an individual input in our novel model architecture named ProtoMask. We conduct experiments on three popular fine-grained classification datasets with a wide set of metrics, providing a detailed overview on explainability characteristics. The comparison with other popular models demonstrates competitive performance and unique explainability features of our model. https://github.com/uos-sis/quanproto

可解释AI原型学习图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。