arXiv:2605.07749cs.CV2026-05

对比三种医学大模型在肾病变分类中的表现,发现传统影像组学仍更优。

Benchmarking Foundation Models for Renal Lesion Stratification in CT

论文配图:Benchmarking Foundation Models for Renal Lesion Stratification in CT
图 1 · 摘自论文原文
  • 用冻结特征探针法比较大模型与手工影像组学、3D ResNet的性能
  • 大模型AUC为0.70-0.77,略低于影像组学的0.88(均p≤0.002)
  • 适合数据稀缺场景,但尚难捕捉细微纹理差异,当前仍以影像组学为最佳

开源医学基础模型(FMs)的快速兴起提出一个实际问题:其预训练表征在临床相关但数据稀缺的分类任务中迁移效果如何?尤其在基于CT的肾病变分类中,提升泛化能力意义重大,因训练数据本就有限。我们通过基准测试评估了三种医学大模型在此任务上的表现。该六类问题涵盖囊肿、透明细胞肾细胞癌及罕见亚型。采用冻结特征探针协议,将大模型嵌入向量与手工影像组学分类器和从零开始训练的3D ResNet-50进行比较。模型在包含2,854个病灶的综合数据集上训练,外部测试集来自癌症成像档案(The Cancer Imaging Archive),共234个病灶。结果显示:第一,大模型性能(AUC 0.70–0.77)与从头训练的ResNet(AUC 0.72)相当,但硬件需求大幅降低,特征提取后仅需数秒即可在CPU上完成推理;第二,传统影像组学基准显著优于所有深度学习方法,达到AUC 0.88(所有p ≤ 0.002)。这表明当前通用型大模型嵌入尚未充分捕捉驱动组织学亚型区分的细粒度纹理与形状异质性。尽管在数据稀缺场景具潜力,医学大模型仍未超越现有成熟模型,影像组学仍是当前肾病变分层的最优方案。

原文摘要 · Abstract (English)

The rapid proliferation of open-source medical foundation models (FMs) raises a practical question: how well do their pre-trained representations transfer to clinically relevant but data-scarce classification tasks? Particularly in CT-based renal lesion classification, a push toward greater generalizability would be meaningful, as the field is constrained by inherently limited training data. We addressed this through a benchmark of three medical FMs on this specific task. This six-class problem spans common entities like cysts and clear cell renal cell carcinoma, alongside rare subtypes. Using a frozen feature-probing protocol, we compared FM embeddings against a handcrafted radiomics classifier and a 3D ResNet-50 trained from scratch. Models were trained on a composite dataset of 2,854 lesions and evaluated on an external test set of 234 lesions from The Cancer Imaging Archive. Our results reveal two key findings. First, FM performance (AUC 0.70-0.77) matched the from-scratch ResNet (AUC 0.72) while drastically reducing hardware demand, requiring only seconds on a CPU after feature extraction. However, the conventional radiomics baseline significantly outperformed all deep learning approaches, achieving an AUC of 0.88 (all p $\leq$ 0.002). This suggests that current generalist FM embeddings do not yet capture the fine-grained texture and shape heterogeneity driving histological subtype discrimination. Despite their potential in data-scarce settings, medical FMs did not surpass established models for renal lesion stratification, leaving radiomics as the current state-of-the-art.

医学大模型肾病变分类影像组学数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。