融合视觉、知识与语言特征,提升乳腺钼靶诊断的准确性与泛化能力。
ViKL: A Mammography Interpretation Framework via Multimodal Aggregation of Visual-knowledge-linguistic Features
- 通过三模态对比学习,融合图像、报告与放射学表现信息。
- 在无病理标签情况下实现性能显著提升,跨数据集迁移能力更强。
- 适合医学影像、多模态学习及临床辅助诊断研究者参考。
乳腺钼靶是乳腺癌诊断的主要影像工具。尽管深度学习在图像解读方面取得进展,但仅依赖视觉特征的方法在跨数据集泛化上仍存在局限。我们提出,整合报告中的语言特征和体现放射学洞察的表现特征,可构建更强大、可解释且泛化的表征。本文发布首个多模态乳腺钼靶数据集MVKL,包含多视角图像、详细表现特征与报告文本。基于此,我们提出ViKL框架,专注于无监督预训练,通过三重对比学习将语言与知识信息与视觉数据融合,实现模态间与模态内特征增强。实验表明:1)结合报告与表现特征,显著提升病理分类性能并促进多模态交互;2)表现特征可引入新颖的困难负样本选择机制;3)多模态特征具备跨数据集迁移能力;4)多模态预训练有效缓解校准偏差,构建高质量表征空间。MVKL数据集与ViKL代码已开源,支持后续研究。
原文摘要 · Abstract (English)
Mammography is the primary imaging tool for breast cancer diagnosis. Despite significant strides in applying deep learning to interpret mammography images, efforts that focus predominantly on visual features often struggle with generalization across datasets. We hypothesize that integrating additional modalities in the radiology practice, notably the linguistic features of reports and manifestation features embodying radiological insights, offers a more powerful, interpretable and generalizable representation. In this paper, we announce MVKL, the first multimodal mammography dataset encompassing multi-view images, detailed manifestations and reports. Based on this dataset, we focus on the challanging task of unsupervised pretraining and propose ViKL, a innovative framework that synergizes Visual, Knowledge, and Linguistic features. This framework relies solely on pairing information without the necessity for pathology labels, which are often challanging to acquire. ViKL employs a triple contrastive learning approach to merge linguistic and knowledge-based insights with visual data, enabling both inter-modality and intra-modality feature enhancement. Our research yields significant findings: 1) Integrating reports and manifestations with unsupervised visual pretraining, ViKL substantially enhances the pathological classification and fosters multimodal interactions. 2) Manifestations can introduce a novel hard negative sample selection mechanism. 3) The multimodal features demonstrate transferability across different datasets. 4) The multimodal pretraining approach curbs miscalibrations and crafts a high-quality representation space. The MVKL dataset and ViKL code are publicly available at https://github.com/wxwxwwxxx/ViKL to support a broad spectrum of future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。