用乳腺影像与临床文本融合,提升癌症早期检测准确率。
Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- 将2D乳腺影像与临床文本通过创新分词模块融合
- 在多国筛查数据集上显著优于单一模态模型
- 适合临床部署,兼顾高分辨率图像与跨人群适用性
乳腺癌是发达国家女性中最常见的恶性肿瘤。通过乳腺钼靶筛查实现早期发现可显著降低死亡率。尽管计算机辅助诊断(CAD)系统对放射科医生有帮助,但现有方法在处理多模态数据的细微解读以及因需病史信息而难以临床落地方面存在关键局限。本研究提出一种新框架,通过创新的分词模块,将二维乳腺影像的视觉特征与易获取的临床元数据及合成放射报告中的结构化文本描述相结合。实验表明,将卷积神经网络(ConvNets)与语言表示策略性融合,在处理高分辨率图像的同时,性能优于基于视觉变换器的模型,并具备实际部署可行性。在多国队列乳腺钼靶数据集上的评估显示,该多模态方法在癌症检测和钙化识别方面均显著优于单模态基线,尤其在跨人群泛化中表现突出。该方法建立了一种新型临床可行的视觉-语言模型CAD系统范式,有效融合影像与患者上下文信息。
原文摘要 · Abstract (English)
Breast cancer remains the most commonly diagnosed malignancy among women in the developed world. Early detection through mammography screening plays a pivotal role in reducing mortality rates. While computer-aided diagnosis (CAD) systems have shown promise in assisting radiologists, existing approaches face critical limitations in clinical deployment - particularly in handling the nuanced interpretation of multi-modal data and feasibility due to the requirement of prior clinical history. This study introduces a novel framework that synergistically combines visual features from 2D mammograms with structured textual descriptors derived from easily accessible clinical metadata and synthesized radiological reports through innovative tokenization modules. Our proposed methods in this study demonstrate that strategic integration of convolutional neural networks (ConvNets) with language representations achieves superior performance to vision transformer-based models while handling high-resolution images and enabling practical deployment across diverse populations. By evaluating it on multi-national cohort screening mammograms, our multi-modal approach achieves superior performance in cancer detection and calcification identification compared to unimodal baselines, with particular improvements. The proposed method establishes a new paradigm for developing clinically viable VLM-based CAD systems that effectively leverage imaging data and contextual patient information through effective fusion mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。