arXiv:2510.14831cs.CV2025-10ICCV被引 19

用合成数据提升肿瘤分割效率,仅需500张真实图像即可达到1500张的效果。

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

  • 利用合成数据配合少量真实数据,突破标注瓶颈
  • 在6个腹部器官上实现+7%到+16%的分割性能提升
  • 适合医学AI研究者和需要高效训练模型的团队

胰腺肿瘤分割的AI模型受限于大规模体素级标注数据的缺乏,而此类数据难以获取且依赖医疗专家。在包含3,000例标注胰腺肿瘤扫描的自有JHH数据集上,我们发现当样本量达到1,500例后,模型性能趋于饱和。通过引入合成数据,仅需500张真实扫描即达到同等性能,表明合成数据可显著加速数据扩展规律。基于此经验,我们构建了AbdomenAtlas 2.0——一个包含10,135例CT扫描、总计15,130个肿瘤实例(六器官:胰腺、肝脏、肾脏、结肠、食道、子宫)和5,893例对照扫描的体素级标注数据集,由23位放射科专家标注。该数据集规模远超现有公开数据集,在分布内测试中分割性能提升7%,分布外测试提升16%。

原文摘要 · Abstract (English)

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after 1,500 scans. With synthetic data, we reached the same performance using only 500 real scans. This finding suggests that synthetic data can steepen data scaling laws, enabling more efficient model training than real data alone. Motivated by these lessons, we created AbdomenAtlas 2.0--a dataset of 10,135 CT scans with a total of 15,130 tumor instances per-voxel manually annotated in six organs (pancreas, liver, kidney, colon, esophagus, and uterus) and 5,893 control scans. Annotated by 23 expert radiologists, it is several orders of magnitude larger than existing public tumor datasets. While we continue expanding the dataset, the current version of AbdomenAtlas 2.0 already provides a strong foundation--based on lessons from the JHH dataset--for training AI to segment tumors in six organs. It achieves notable improvements over public datasets, with a +7% DSC gain on in-distribution tests and +16% on out-of-distribution tests.

肿瘤分割合成数据医学影像数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。