融合全局与局部特征,提升远端肌病影像诊断准确率与可解释性。
Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
- 通过注意力门机制融合双模型特征,兼顾上下文与细节信息。
- 在公开与自建数据集上均达高分类准确率,生成临床相关显著图。
- 专家评估发现解释性仍有不足,需结合医生反馈优化可解释性。
远端肌病是一组遗传异质性骨骼肌疾病,临床表现多样,给放射科诊断带来挑战。为此,我们提出一种新型多模态注意力感知融合架构,结合两个深度学习模型提取的特征:一个捕捉全局上下文信息,另一个聚焦局部细节,二者互补。独特之处在于,通过注意力门机制融合特征,同时提升预测性能与可解释性。该方法在BUSI基准数据集和自建远端肌病数据集上均取得高分类准确率,并生成具有临床意义的显著性热力图,支持透明化医疗决策。我们通过功能基度量、参考掩码一致性评分及增量删除分析,以及七位专家放射科医师的应用验证,严格评估了可解释性。尽管融合策略优于单流及其它融合方式,定量与定性评估仍显示解释结果在解剖特异性与临床实用性方面存在持续差距,凸显需要更丰富、情境感知的可解释方法,以及人机协同反馈以满足真实诊疗场景中临床医生的期待。
原文摘要 · Abstract (English)
Distal myopathy represents a genetically heterogeneous group of skeletal muscle disorders with broad clinical manifestations, posing diagnostic challenges in radiology. To address this, we propose a novel multimodal attention-aware fusion architecture that combines features extracted from two distinct deep learning models, one capturing global contextual information and the other focusing on local details, representing complementary aspects of the input data. Uniquely, our approach integrates these features through an attention gate mechanism, enhancing both predictive performance and interpretability. Our method achieves a high classification accuracy on the BUSI benchmark and a proprietary distal myopathy dataset, while also generating clinically relevant saliency maps that support transparent decision-making in medical diagnosis. We rigorously evaluated interpretability through (1) functionally grounded metrics, coherence scoring against reference masks and incremental deletion analysis, and (2) application-grounded validation with seven expert radiologists. While our fusion strategy boosts predictive performance relative to single-stream and alternative fusion strategies, both quantitative and qualitative evaluations reveal persistent gaps in anatomical specificity and clinical usefulness of the interpretability. These findings highlight the need for richer, context-aware interpretability methods and human-in-the-loop feedback to meet clinicians' expectations in real-world diagnostic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。