27篇结肠镜息肉分割论文暴露出评估标准混乱问题,影响结果可信度。
Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

- 3大问题:漏报豪斯多夫距离、训练测试划分不一致、无统计显著性检验
- 同一数据集上不同论文的骰子系数无法比较,最佳模型随指标变化
- 提出可落地的息肉分割报告清单,助力规范评估
结肠镜息肉分割的研究进展通常依赖于少数公开基准上的排行榜对比。我们指出,这种表面进展难以验证:对2015至2026年间发表的27篇论文进行系统审计,发现三大结构性问题。首先,25篇论文遗漏了具有临床意义的豪斯多夫距离(Hausdorff distance)——该指标对检测扁平或小型息肉至关重要,且在放射治疗分割中为标准。其次,至少五个不兼容的训练/测试划分协议共存于相同两个数据集(Kvasir-SEG 和 CVC-ClinicDB)上,导致公布的骰子系数不可比。第三,26篇论文在无任何统计显著性检验的情况下做出性能宣称。令人震惊的是,四篇发表于《Metrics Reloaded》框架(Maier-Hein et al., Nature Methods 2024)之后的论文仍延续这些问题,表明通用评估指导尚未深入该子领域。为验证问题严重性,我们在三种受控协议下使用统一评分器重新评估五种代表性模型,发现报告指标掩盖了严重的边界和召回失败,最优模型随指标变化,近似排名在随机划分下反转。为此,我们提出轻量级的五点「息肉分割报告检查清单」(PSRC),以推动领域内评估规范化。
原文摘要 · Abstract (English)
Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this apparent progress is difficult to verify: a systematic audit of \textbf{27 papers} published between 2015 and 2026 reveals three structural problems in how the community evaluates models. \textbf{First}, 25 of 27 papers \textit{omit the Hausdorff distance}. Hausdorff distance is a boundary-accuracy metric with direct clinical relevance for detecting flat or small polyps, and is a standard in radiotherapy segmentation. \textbf{Second}, at least five \textit{incompatible train/test split protocols} co-exist across papers reporting results on the same two datasets (Kvasir-SEG and CVC-ClinicDB), making published Dice scores non-comparable even when they appear in the same leaderboard column. \textbf{Third}, 26 of 27 papers make \textit{performance claims without any statistical significance test}. Strikingly, four papers published \emph{after} the Metrics Reloaded framework~\cite{metricsreloaded2024} (Maier-Hein et al., \textit{Nature Methods} 2024) perpetuate these same problems, suggesting that general-purpose metric guidance has not yet reached the colonoscopy sub-community. To show these problems are not merely cosmetic, we re-evaluate five representative models under three controlled protocols with a single uniform scorer, and find that the reported metric conceals large boundary and recall failures, that the ``best'' model changes with the metric, and that near-tied rankings reverse across random splits. We propose a five-point \textbf{Polyp Segmentation Reporting Checklist}~(PSRC) as a lightweight, domain-adapted corrective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。