arXiv:2505.17367cs.CVcs.AI2025-05

提出可解释的视觉马尔可夫架构,提升多器官医学图像分类准确率与可信度。

EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion

  • 采用多路径设计,融合密集网、U-Net与视觉马尔可夫模块,动态集成特征。
  • 在9类多器官数据集上达到99.75%测试准确率,显著优于现有方法。
  • 支持多种可解释性机制,适合临床可信AI诊断场景使用。

医学图像分类对临床决策至关重要,但准确性、可解释性和泛化能力仍具挑战。本文提出EVM-Fusion,一种基于新型神经算法融合(NAF)机制的可解释视觉马尔可夫架构,用于多器官医学图像分类。该模型采用多路径设计:由密集网(DenseNet)和U-Net构建的路径,结合视觉马尔可夫(Vim)模块,与传统特征路径并行运行。通过两阶段融合过程——跨模态注意力与迭代NAF块——实现特征的动态整合,其中NAF块可学习自适应融合策略。内在可解释性通过路径特异性空间注意力、Vim Δ值图、传统特征SE注意力及跨模态注意力权重实现。在包含9类器官的多样化医学图像数据集上的实验表明,EVM-Fusion在测试中达到99.75%的准确率,并提供多维度决策洞察,展现出在医疗诊断中构建可信AI的巨大潜力。

原文摘要 · Abstract (English)

Medical image classification is critical for clinical decision-making, yet demands for accuracy, interpretability, and generalizability remain challenging. This paper introduces EVM-Fusion, an Explainable Vision Mamba architecture featuring a novel Neural Algorithmic Fusion (NAF) mechanism for multi-organ medical image classification. EVM-Fusion leverages a multipath design, where DenseNet and U-Net based pathways, enhanced by Vision Mamba (Vim) modules, operate in parallel with a traditional feature pathway. These diverse features are dynamically integrated via a two-stage fusion process: cross-modal attention followed by the iterative NAF block, which learns an adaptive fusion algorithm. Intrinsic explainability is embedded through path-specific spatial attention, Vim Δ-value maps, traditional feature SE-attention, and cross-modal attention weights. Experiments on a diverse 9-class multi-organ medical image dataset demonstrate EVM-Fusion's strong classification performance, achieving 99.75% test accuracy and provide multi-faceted insights into its decision-making process, highlighting its potential for trustworthy AI in medical diagnostics.

可解释AI医学图像视觉马尔可夫多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。