用合成报告训练多视角乳腺影像模型,提升癌症诊断与风险预测性能。
MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- 通过跨模态自监督学习融合多视角影像与伪报告,构建联合视觉语言表征。
- 在三个分类任务中均达到顶尖效果,且仅用合成文本即超越全监督基线。
- 适合医学影像研究者与临床辅助诊断系统开发者使用。
大型标注数据集对于训练鲁棒的乳腺癌检测或风险预测计算机辅助诊断(CAD)模型至关重要,但精细标注成本高、耗时长。视觉-语言模型(如CLIP)在大规模图像-文本对上预训练,可提升医学影像任务的鲁棒性与数据效率。本文提出一种新型多视角乳腺摄影与语言模型(MV-MLM),基于配对的乳腺钼靶图像与合成放射科报告进行训练。该模型利用多视角监督和跨模态自监督机制,从大量放射科数据中学习丰富表征,包含多个视角及其对应的伪放射科报告。我们设计了一种新的联合视觉-文本学习策略,以增强不同数据类型和任务下的泛化能力与准确率,从而区分乳腺组织或癌变特征(钙化、肿块),并据此理解影像内容、预测癌症风险。在私有及公开数据集上的评估表明,所提模型在三项分类任务中表现最优:(1) 恶性程度分类,(2) 亚型分类,(3) 基于图像的癌症风险预测。此外,该模型展现出强数据效率,在仅使用合成文本报告且无需真实放射科报告的情况下,优于现有全监督或VLM基线。
原文摘要 · Abstract (English)
Large annotated datasets are essential for training robust Computer-Aided Diagnosis (CAD) models for breast cancer detection or risk prediction. However, acquiring such datasets with fine-detailed annotation is both costly and time-consuming. Vision-Language Models (VLMs), such as CLIP, which are pre-trained on large image-text pairs, offer a promising solution by enhancing robustness and data efficiency in medical imaging tasks. This paper introduces a novel Multi-View Mammography and Language Model for breast cancer classification and risk prediction, trained on a dataset of paired mammogram images and synthetic radiology reports. Our MV-MLM leverages multi-view supervision to learn rich representations from extensive radiology data by employing cross-modal self-supervision across image-text pairs. This includes multiple views and the corresponding pseudo-radiology reports. We propose a novel joint visual-textual learning strategy to enhance generalization and accuracy performance over different data types and tasks to distinguish breast tissues or cancer characteristics(calcification, mass) and utilize these patterns to understand mammography images and predict cancer risk. We evaluated our method on both private and publicly available datasets, demonstrating that the proposed model achieves state-of-the-art performance in three classification tasks: (1) malignancy classification, (2) subtype classification, and (3) image-based cancer risk prediction. Furthermore, the model exhibits strong data efficiency, outperforming existing fully supervised or VLM baselines while trained on synthetic text reports and without the need for actual radiology reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。