arXiv:2605.05775cs.CVcs.AI2026-05被引 12

挑战多中心多示踪剂PET/CT自动病灶分割,突破数据泛化瓶颈。

The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT $\unicode{x2013}$ Multitracer Multicenter Generalization

论文配图:The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT $\unicode{x2013}$ Multitracer Multicenter Generalization
图 1 · 摘自论文原文
  • 基于3D nnU-Net融合PET/CT通道,实现跨中心多示踪剂分割
  • 最佳模型平均骰率0.66,假阴性体积减少5mL,优于基线8%
  • 揭示系统性高估是泛化难题主因,数据差异影响远超算法选择

我们报告了2024年MICCAI第三届autoPET挑战赛的设计与结果,该挑战在组合泛化设置下评估全身影像中肿瘤病灶的自动化分割。训练数据包含图宾根大学医院1,014例[18F]-FDG PET/CT与慕尼黑路德维希马克西米利安大学医院597例[18F]/[68Ga]-PSMA PET/CT,构成迄今最大公开可获取的PSMA PET/CT标注数据集。测试集共200例,涵盖四种示踪剂-中心组合,其中两种为未见的组合配对。另设数据驱动奖类别,限制参赛者使用固定基线模型以分离数据处理策略贡献。共17支队伍提交27个算法,主要为基于nnU-Net的3D网络,采用PET/CT通道拼接。最优算法在所有四类测试条件下均取得平均骰率(DSC)0.66、假阴性体积(FNV)3.18 mL、假阳性体积(FPV)2.78 mL,相比基线提升8%骰率,假阴性体积减少5 mL。排名在自助抽样和替代评分方案下保持稳定。深入分析表明:(1) 同域多示踪剂分割已足够接近放射科医生一致性;(2) 未知示踪剂-中心组合的泛化仍为开放问题,主因是系统性体积高估;(3) 患者间异质性与病例难度对性能影响远大于顶尖团队间的算法选择差异。

原文摘要 · Abstract (English)

We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body PET/CT under a compositional generalization setting. Training data comprised 1,014 [18F]-FDG PET/CT studies from the University Hospital Tübingen and 597 [18F]/[68Ga]-PSMA PET/CT studies from the LMU University Hospital Munich, constituting the largest publicly available annotated PSMA PET/CT dataset to date. The held-out test set of 200 studies covered four tracer-center combinations, two of which represented unseen compositional pairings. A complementary data-centric award category isolated the contribution of data handling strategies by restricting participants to a fixed baseline model. Seventeen teams submitted 27 algorithms, predominantly nnU-Net-based 3D networks with PET/CT channel concatenation. The top-ranked algorithm achieved a mean DSC of 0.66, FNV of 3.18 mL, and FPV of 2.78 mL across all four test conditions, improving DSC by 8% and reducing the false-negative volume by 5 mL relative to the provided baseline. Ranking was stable across bootstrap resampling and alternative ranking schemes for the top tier. Beyond the benchmark, we provide an in-depth analysis of segmentation performance at the patient and lesion level. Three main conclusions can be drawn: (1) in-domain multitracer PET/CT segmentation is sufficient and probably approaching reader agreement; (2) compositional generalization to unseen tracer-center combinations remains an open problem mainly driven by systematic volume overestimation; (3) heterogeneity and case difficulty drive performance variation substantially more than the choice of algorithm among top-ranked teams.

PET/CT病灶分割多中心泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。