对比基础模型与影像组学在肺CT中的表现,发现不同任务需选不同方案。
Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices

- 分阶段设计避免小样本过拟合,聚焦特征提取器、分类头和分割选择的影响
- 肿瘤体积与分期靠分割,生存、病理和年龄预测则取决于分类器选择
- 推荐用Curia+肿瘤分割+CatBoost作为通用方案,但特定任务定制更优
影像组学是基于CT的肺癌表型分析的主流方法,但与基础模型的对比常未能分离特征提取器、分类头和分割策略的贡献,也缺乏跨队列鲁棒性验证。本研究在五个任务上基准测试了五种特征提取器(Curia、Curia-2、DINOv3、Radiomics2D、Radiomics3D)、七种分类头(TabPFN、TabICL、XGBoost、CatBoost、Random Forest、logistic regression、Ridge)和三种分割方案:肿瘤体积与分期分类、2年生存率预测、组织学分类及年龄预测。模型在LUNG1(n=338)上训练,在内部测试集(n=84)和外部LUNG2队列(n=211)上评估,以最差跨队列表现作为主要指标。结果显示,任务依赖主导因素:分割影响体积与分期,分类器选择决定生存、组织学和年龄预测。影像组学在肿瘤体积、分期和生存预测中表现竞争性(部分源于标签衍生效应);Curia系列在生存预测中达到相近峰值性能;DINOv3整体略逊。图像块与切片聚合对结果影响可忽略。建议采用Curia+肿瘤分割+CatBoost作为安全默认方案,在三大临床任务中取得最佳平均排名,但任务特异性选择始终更优。若无肿瘤标注,可用Curia-2+肺部分割+逻辑回归作为替代方案。所有流程均采用双阶段设计,适用于小样本场景,避免端到端微调导致过拟合。
原文摘要 · Abstract (English)
Radiomics is the established approach for CT-based lung cancer phenotyping, yet comparisons with foundation models rarely isolate contributions of feature extractor, classification head, and segmentation choice, or test cross-cohort robustness. We benchmark five feature extractors (Curia, Curia-2, DINOv3, Radiomics2D, Radiomics3D), seven classification heads (TabPFN, TabICL, XGBoost, CatBoost, Random Forest, logistic regression, Ridge), and three segmentation regimes on five tasks: tumor volume and stage classification, 2-year survival prediction, histology classification, and age prediction. Models are trained on LUNG1 (n=338) and evaluated on an internal test set (n=84) and the external LUNG2 cohort (n=211), with worst-case cross-cohort performance as the primary metric. The dominant design factor is task-dependent: segmentation drives volume and stage classification, while classifier choice drives survival, histology, and age prediction. Radiomics is competitive for tumor volume, tumor stage and survival (partly due to label-derivation effects for the former); Curia variants reach comparable peak scores for survival; DINOv3 falls slightly short across tasks. Patch and slice aggregation have negligible impact. We recommend Curia with tumor segmentation and a CatBoost head as a safe default, achieving the best mean rank across the three primary clinical tasks, though task-specific selection consistently outperforms any cross-task default. When tumor delineations are unavailable, Curia-2 with lung segmentation and logistic regression offers a competitive alternative. All pipelines use a two-stage design suited to small cohort sizes where end-to-end fine-tuning would risk overfitting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。