arXiv:2507.05582eess.IVcs.CV2025-07中稿 · MICCAI 2025被引 12

用放射科报告指导模型分割肿瘤,提升小样本下的准确率

Learning Segmentation from Radiology Reports

  • 将报告转化为体素级监督信号,弥补标注掩码不足
  • 在50个样本下F1分数提升16%,1.7千样本时也显著增益
  • 适合医疗影像数据少、但报告多的场景

CT扫描中的肿瘤分割对诊断、手术和预后至关重要,但分割掩码稀缺,因其制作需时间和专业技能。公共腹部CT数据集仅有几十到上千个肿瘤掩码,而医院存有数十万张带报告的肿瘤CT。因此,利用报告提升分割性能是关键。本文提出报告监督损失(R-Super),将放射科报告转换为体素级监督信号用于肿瘤分割。我们构建了包含6,718对CT-报告的数据集(来自UCSF医院),并融合公开的CT-掩码数据集(AbdomenAtlas 2.0)。使用R-Super结合掩码与报告训练,在内部和外部验证中均显著提升分割效果:相较于仅用掩码训练,F1分数最高提升16%。该方法在极少数掩码(如50个)和较多掩码(如1.7千个)情况下均表现优异,通过利用现成的报告补充稀疏标注,有效提升AI性能。

原文摘要 · Abstract (English)

Tumor segmentation in CT scans is key for diagnosis, surgery, and prognosis, yet segmentation masks are scarce because their creation requires time and expertise. Public abdominal CT datasets have from dozens to a couple thousand tumor masks, but hospitals have hundreds of thousands of tumor CTs with radiology reports. Thus, leveraging reports to improve segmentation is key for scaling. In this paper, we propose a report-supervision loss (R-Super) that converts radiology reports into voxel-wise supervision for tumor segmentation AI. We created a dataset with 6,718 CT-Report pairs (from the UCSF Hospital), and merged it with public CT-Mask datasets (from AbdomenAtlas 2.0). We used our R-Super to train with these masks and reports, and strongly improved tumor segmentation in internal and external validation--F1 Score increased by up to 16% with respect to training with masks only. By leveraging readily available radiology reports to supplement scarce segmentation masks, R-Super strongly improves AI performance both when very few training masks are available (e.g., 50), and when many masks were available (e.g., 1.7K). Project: https://github.com/MrGiovanni/R-Super

医学图像弱监督报告生成分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。