arXiv:2606.23172eess.IV2026-06

用MRI预测胶质瘤基因突变,图像与表格模型各有优劣

A Benchmark of (MRI-) Foundation Models to Predict IDH Mutational Status in Glioma

论文配图:A Benchmark of (MRI-) Foundation Models to Predict IDH Mutational Status in Glioma
图 1 · 摘自论文原文
  • 比较了四种图像基础模型和表型模型在基因预测中的表现
  • 表格模型在多数数据集上表现最佳,图像模型在外部数据上更具潜力
  • 模型性能受数据分布影响大,需关注临床实际场景差异

从常规磁共振成像(MRI)非侵入性预测胶质瘤分子状态已展现出良好性能,但受限于影像-基因组匹配数据集规模小,模型泛化能力仍存挑战。基础模型可能缓解此瓶颈,但需全面基准评估不同架构、预训练领域和目标的影响。针对从FLAIR和增强T1 MRI预测异柠檬酸脱氢酶(IDH)突变的任务,我们对比了四种基于图像的基础模型(BrainIAC、MRI-CORE、BiomedCLIP、BrainDINO)与基于放射组学的TabPFN及逻辑回归基线。在四个公开成人胶质瘤队列和一个外部治疗后队列中评估了预测性能与校准性。同队列下,TabPFN达到0.92(0.03)AUROC和0.74(0.17)AUPRC(均值±标准差),优于或相当所有视觉编码器;其中BiomedCLIP表现最佳(0.85(0.08)AUROC),BrainDINO亦表现良好(0.82(0.09)AUROC),而专用MRI编码器(BrainIAC、MRI-CORE)持续表现较差。跨队列迁移中出现中等程度的AUROC下降,但AUPRC对患病率变化更敏感。在外部队列中,BiomedCLIP取得最高AUROC(0.74(0.07)),而TabPFN提供更好校准性(期望校准误差0.07(0.01))。结果表明,表示模态与评估上下文显著影响基础模型在MRI分子预测中的表现。基于放射组学特征的表格基础模型构成强且校准良好的基线,而图像基础模型在临床分布差异明显时可能具有互补价值。

原文摘要 · Abstract (English)

Non-invasive prediction of glioma molecular status from routine magnetic resonance imaging (MRI) has shown promising performance, but model generalization remains challenging given small-scale matched imaging-genomic datasets. Foundation models may address this bottleneck, but a comprehensive benchmark is needed to establish the impact of diverse architectures, pre-training domains, and objectives. Given the use case of isocitrate dehydrogenase (IDH) mutation prediction from FLAIR and post-contrast T1 MRIs, we compared four image-based foundation models, BrainIAC, MRI-CORE, BiomedCLIP, and BrainDINO, against radiomics-based TabPFN and logistic regression baselines. Prediction performance and calibration were assessed across four public adult glioma cohorts and an external post-treatment cohort. Within-cohort, TabPFN matched or outperformed all visual encoders, achieving 0.92 (0.03) AUROC and 0.74 (0.17) AUPRC (mean (SD) across all datasets). Among visual encoders, BiomedCLIP performed best (0.85 (0.08) AUROC), with BrainDINO competitive (0.82 (0.09) AUROC), while MRI-specific encoders (BrainIAC, MRI-CORE) consistently underperformed. Cross-cohort transfer showed moderate AUROC degradation but stronger AUPRC sensitivity to prevalence shifts. On the external cohort, BiomedCLIP achieved the highest AUROC (0.74 (0.07)), whereas TabPFN provided superior calibration (Expected Calibration Error 0.07 (0.01)). These results indicate that representation modality and evaluation context critically influence foundation-model performance in MRI-based molecular prediction. Tabular foundation models on radiomic features provide a strong, well-calibrated baseline, while image foundation models may offer complementary value under clinically distinct distribution shifts. Code available at https://github.com/nathanhollet/idh-status-prediction

医学影像胶质瘤基础模型基因预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。