arXiv:2508.18612eess.IVcs.LG2025-08被引 1

跨癌种测试3D nnU-Net在PET-CT肿瘤分割中的泛化能力,发现数据多样性比模型复杂度更重要。

Stress-testing cross-cancer generalizability of 3D nnU-Net for PET-CT tumor segmentation: multi-cohort evaluation with novel oesophageal and lung cancer datasets

  • 融合多中心、多癌种数据训练提升模型泛化性
  • 联合训练组在三组数据上平均DSC均超50%
  • 强调真实临床场景下数据多样性对模型鲁棒性的关键作用

深度学习在临床PET-CT肿瘤分割中需具备强泛化能力,因解剖部位、扫描设备和患者群体差异大。本研究首次在PET-CT上开展跨癌种的nnU-Net评估,引入两个全新专家标注的全身数据集:279例食管癌(澳大利亚队列)和54例肺癌(印度队列)。这些队列补充了公开的AutoPET数据集,实现跨域性能的系统性压力测试。我们以三种范式训练3D nnUNet:仅食管癌训练、仅AutoPET训练、联合训练。在测试集上,仅食管癌模型在本域表现最佳(平均DSC 57.8),但在外部印度肺癌队列上失败(平均DSC <3.4),表明严重过拟合。仅AutoPET模型泛化更广(AutoPET平均DSC 63.5,印度肺癌队列51.6),但在澳大利亚食管癌队列上表现差(平均DSC 26.7)。联合训练方案结果最均衡(肺癌平均DSC 52.9,食管癌40.7,AutoPET 60.9),减少边界误差,提升跨队列鲁棒性。结果表明,数据多样性——尤其是多人群、多中心、多癌种整合——是实现稳健泛化的决定性因素,远超模型架构创新。该工作建立基于人口统计的跨癌种深度学习分割评估框架,强调数据多样性而非模型复杂度才是临床可靠分割的基础。

原文摘要 · Abstract (English)

Robust generalization is essential for deploying deep learning based tumor segmentation in clinical PET-CT workflows, where anatomical sites, scanners, and patient populations vary widely. This study presents the first cross cancer evaluation of nnU-Net on PET-CT, introducing two novel, expert-annotated whole-body datasets. 279 patients with oesophageal cancer (Australian cohort) and 54 with lung cancer (Indian cohort). These cohorts complement the public AutoPET dataset and enable systematic stress-testing of cross domain performance. We trained and tested 3D nnUNet models under three paradigms. Target only (oesophageal), public only (AutoPET), and combined training. For the tested sets, the oesophageal only model achieved the best in-domain accuracy (mean DSC, 57.8) but failed on external Indian lung cohort (mean DSC less than 3.4), indicating severe overfitting. The public only model generalized more broadly (mean DSC, 63.5 on AutoPET, 51.6 on Indian lung cohort) but underperformed in oesophageal Australian cohort (mean DSC, 26.7). The combined approach provided the most balanced results (mean DSC, lung (52.9), oesophageal (40.7), AutoPET (60.9)), reducing boundary errors and improving robustness across all cohorts. These findings demonstrate that dataset diversity, particularly multi demographic, multi center and multi cancer integration, outweighs architectural novelty as the key driver of robust generalization. This work presents the demography based cross cancer deep learning segmentation evaluation and highlights dataset diversity, rather than model complexity, as the foundation for clinically robust segmentation.

肿瘤分割PET-CT泛化性数据多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。