用双注意力+概念图提升医学诊断可解释性
DCG-Net: Dual Cross-Attention with Concept-Value Graph Reasoning for Interpretable Medical Diagnosis
- 双交叉注意力对齐图像与概念原型,定位关键证据区域
- 基于互信息先验的概念图建模临床概念间依赖关系
- 在血细胞和皮肤病变诊断中兼具高精度与可解释性
深度学习在医学图像分析中表现优异,但其决策过程难以解释。概念瓶颈模型(CBMs)通过人类可理解的临床概念结构化预测,部分缓解此问题,但现有方法常忽略概念间的上下文依赖。为此,我们提出端到端可解释框架DCG-Net,融合多模态对齐与结构化概念推理。DCG-Net引入双交叉注意力模块,以双向注意力替代余弦相似度匹配,实现视觉标记与标准化文本概念-值原型之间的交互,支持空间局部化证据归因。为捕捉临床概念固有的关系结构,我们构建了基于正点互信息先验的参数化概念图,并通过稀疏控制的消息传递进行优化。该设计符合临床领域知识。在白血球形态与皮肤病变诊断任务上的实验表明,DCG-Net在取得当前最优分类性能的同时,生成了具有临床意义的诊断解释。
原文摘要 · Abstract (English)
Deep learning models have achieved strong performance in medical image analysis, but their internal decision processes remain difficult to interpret. Concept Bottleneck Models (CBMs) partially address this limitation by structuring predictions through human-interpretable clinical concepts. However, existing CBMs typically overlook the contextual dependencies among concepts. To address these issues, we propose an end-to-end interpretable framework \emph{DCG-Net} that integrates multimodal alignment with structured concept reasoning. DCG-Net introduces a Dual Cross-Attention module that replaces cosine similarity matching with bidirectional attention between visual tokens and canonicalized textual concept-value prototypes, enabling spatially localized evidence attribution. To capture the relational structure inherent to clinical concepts, we develop a Parametric Concept Graph initialized with Positive Pointwise Mutual Information priors and refined through sparsity-controlled message passing. This formulation models inter-concept dependencies in a manner consistent with clinical domain knowledge. Experiments on white blood cell morphology and skin lesion diagnosis demonstrate that DCG-Net achieves state-of-the-art classification performance while producing clinically interpretable diagnostic explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。