融合影像与病历的几何多模态模型,提升前列腺癌分类准确率
A Geometric Multimodal Foundation Model Integrating Bp-MRI and Clinical Reports in Prostate Cancer Classification
- 用几何结构建模医学影像与文本信息,学习联合表征
- 仅用10%数据即达AUC-PR 90.67,优于传统方法8.3个百分点
- 适合临床辅助诊断场景,尤其数据稀缺时表现稳健
前列腺癌(PCa)是全球男性中最常见的癌症之一。双参数MRI(bp-MRI)和临床变量对PCa的识别及治疗决策至关重要,但现有方法高度依赖专家主观判断。多数现有的计算机辅助诊断模型仅关注影像数据,忽视临床背景,且受限于数据稀缺,难以学习鲁棒表征。本文提出一种几何多模态基础模型(MFM-Geom),从bp-MRI和临床报告中联合学习表征,编码视觉发现与临床变量上下文信息。分类头利用对称正定(SPD)矩阵与黎曼深度学习,整合多模态表征。在仅使用10%训练数据的情况下,MFM-Geom相较于基线类标记嵌入方法提升8.3%(AUC-PR达90.67)。在外部数据集上验证其泛化能力,微调后仍达到AUC-PR 90.6。
原文摘要 · Abstract (English)
Prostate cancer (PCa) is one of the most common cancers in men worldwide. Bi-parametric MRI (bp-MRI) and clinical variables are crucial for PCa identification and improving treatment decisions. However, this process is subjective to expert interpretations. Furthermore, most existing computer-aided diagnosis methods focus on imaging-based models, overlooking the clinical context and suffering from data scarcity, limiting their ability to learn robust representations. We propose a geometric multimodal Foundation Model (FM), named MFM-Geom, that learns representations from bp-MRI and clinical reports, encoding visual findings and information from the context of clinical variables. In the representations classification head, the approach leverages symmetric positive definite (SPD) matrices and Riemannian deep learning to integrate imaging-text representations from a biomedical multimodal FM. Using 10% of the training data, MFM-Geom outperformed baseline class token embedding-based classification (+8.3%, AUC-PR of 90.67). Generalization on external dataset confirmed the robustness of fine-tuning biomedical FM, achieving an AUC-PR of 90.6.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。