双分支多模态框架提升医学图像异常检测准确率
DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection

- 构建文本-图像与视觉双分支,融合多模态信息识别异常样本
- 在多个内镜数据集上性能提升最高达24.84%
- 适合医疗AI可靠性增强、跨模态异常检测研究者
复杂多变的真实临床环境对深度学习系统的可靠性提出更高要求。当模型遇到偏离训练分布的数据(如未见疾病病例)时,分布外(OOD)检测对提升模型可靠性与泛化能力至关重要。然而现有方法通常依赖单一视觉模态或仅基于图像-文本匹配,未能充分利用多模态信息。为此,我们提出一种新颖的双分支多模态框架,包含文本-图像分支与视觉分支,通过两个互补分支充分挖掘多模态表征以识别分布外样本。训练完成后,分别计算文本-图像分支得分$S_t$与视觉分支得分$S_v$,并融合得到最终的OOD得分$S$,再与阈值比较实现检测。在多个公开可用的内镜图像数据集上的全面实验表明,所提框架对多种骨干网络均具鲁棒性,且在OOD检测上相比当前最优方法性能提升最高达24.84%。
原文摘要 · Abstract (English)
The complex and dynamic real-world clinical environment demands reliable deep learning (DL) systems. Out-of-distribution (OOD) detection plays a critical role in enhancing the reliability and generalizability of DL models when encountering data that deviate from the training distribution, such as unseen disease cases. However, existing OOD detection methods typically rely either on a single visual modality or solely on image-text matching, failing to fully leverage multimodal information. To overcome the challenge, we propose a novel dual-branch multimodal framework by introducing a text-image branch and a vision branch. Our framework fully exploits multimodal representations to identify OOD samples through these two complementary branches. After training, we compute scores from the text-image branch ($S_t$) and vision branch ($S_v$), and integrate them to obtain the final OOD score $S$ that is compared with a threshold for OOD detection. Comprehensive experiments on publicly available endoscopic image datasets demonstrate that our proposed framework is robust across diverse backbones and improves state-of-the-art performance in OOD detection by up to 24.84%
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。