arXiv:2608.24486eess.IVcs.CV2026-08

标注质量比模型改进对肺栓塞分割效果影响更大

Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

论文配图:Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation
图 1 · 摘自论文原文
  • 用人工复核标注数据,发现标注差异可显著提升分割性能
  • 标注改进使DSC提升0.14~0.19,远超模型调整的0.028
  • 公开了人类参照评估框架,适合医学图像分割研究者使用

目的:量化评估标注对肺栓塞(PE)分割性能的影响,相比模型训练变化,建立人类参照评估框架。方法:回顾性筛选166例体素级标注的CT肺动脉造影病例(CADPE n=91,FUMPE n=35,READ n=40),最终纳入149例。一名主评者按标准流程标注PE,资深胸科放射科医师复核并修正所有分割结果。三家中心另三位评者对15例子集进行标注。通过对比两个预训练nnU-Net模型(nnU-Net-A、nnU-Net-B)在原始与优化标注下的表现,评估标注效应;固定标注,比较同一架构在不同数据集组合训练下的性能,评估模型效应。基准模型(nnPE)采用留一数据集外与五折交叉验证训练。分析四类指标,采用成对威尔科克森符号秩检验、贝叶斯校正及自助法95%置信区间。结果:仅变更标注时,nnU-Net-A和nnU-Net-B的平均DSC分别提升0.143(0.122–0.166)和0.188(0.163–0.213)(均P < .001);而更换训练数据集仅使DSC提升0.028。标注效应在CADPE和FUMPE上超过模型效应,在READ上为0.045。重标注后三组数据的像素强度标准差均显著下降(均P < .001)。nnPE在交叉验证中达到DSC 0.72 ± 0.22,但在52次配对比较中均低于四位标注者(校正后P < .05)。结论:评估标注对测量的肺栓塞分割性能影响至少与模型选择相当。本研究建立了公开的人类参照评估框架,可供未来研究使用。

原文摘要 · Abstract (English)

Purpose: To quantify how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, and to establish a human-referenced framework. Materials and Methods: This retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n=91), FUMPE (n=35), and READ (n=40); 149 were included. A primary rater annotated PE by protocol, and a senior thoracic radiologist reviewed and revised all segmentations. Three additional raters at three centers annotated a 15-case subset. The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A, nnU-Net-B) against original and refined annotations. The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed. The benchmark model (nnPE) was trained with leave-one-dataset-out and pooled five-fold cross-validation. Four metric categories were analyzed with case-paired Wilcoxon signed-rank tests, Benjamini-Hochberg correction, and bootstrap 95% CIs. Results: Changing only the annotation increased mean DSC by 0.143 (0.122-0.166) for nnU-Net-A and 0.188 (0.163-0.213) for nnU-Net-B (both P < .001), whereas changing training-dataset composition changed DSC by 0.028. The label effect exceeded the model effect on CADPE and FUMPE and was 0.045 on READ. Within-mask attenuation SD fell in all three datasets after re-annotation (all P < .001). nnPE reached DSC 0.72 +/- 0.22 on pooled cross-validation but scored below all four annotators across 52 paired comparisons (all corrected P < .05). Conclusion: Evaluation annotations affected measured PE segmentation performance at least as much as model training choices. A human-referenced evaluation framework is publicly available for future study.

医学影像分割评估标注质量深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。