arXiv:2608.30835cs.CVcs.AI2026-08

提出可复现的病理图像伪影检测评估协议,区分真实效果与统计噪音。

Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis

  • 构建四重变异分析框架,验证结果是否真正可靠。
  • 24张幻灯片中70%标注像素集中在4张上,有效样本仅6.2张。
  • 评估协议成本低,适合所有小样本病理研究复现验证。

质量控制是全幻灯片图像分析的前提,但现有基准存在四个问题:独立幻灯片少、标注集中于少数几张、合并比率指标无闭式标准误、沿用单一训练/测试划分。本文提出可靠性评估协议,量化测试集抽样、训练随机性、划分构成和未记录预处理四种变异来源;只有通过全部四项检验的结果才可报告。在对一篇已发表扩散模型伪影检测器的独立重建中,该方法确认辅助对比损失将整体F1从0.673提升至0.688,并在第二随机种子下重现(+0.0156, p=0.031;+0.0190, p=0.005),但其作用机制针对笔迹而非所声称的伪影类型。相比监督基线及其他设计变体的差异均落在评估不确定性范围内。24张幻灯片中4张贡献了70%标注像素,有效样本量为6.2,继承划分位于第7百分位。未报告的组织限制步骤排除了41.4%的模糊标注,而仅排除2.6%的气泡,此筛选本身即与模糊混淆。结论:小样本基准支持的结论远弱于当前报告程度。四项检查成本低廉,可伴随任何此类资源评估,有效分离可复现效应与评估无法分辨的差异。

原文摘要 · Abstract (English)

Background and Objective: Quality control is a prerequisite for whole-slide image analysis, yet the benchmarks on which quality-control methods are compared share four properties that make their reported differences hard to interpret: few independent slides, annotation concentrated in a minority of them, pooled ratio metrics with no closed-form standard error, and a single inherited train/test partition. We propose a reliability protocol for such benchmarks. Methods: The protocol quantifies four sources of variability - test-set sampling, training stochasticity, partition composition, and undocumented preprocessing - a claim is reportable only if it survives all four; three of the four cost minutes of compute. We apply it to an independent reconstruction of a published diffusion-based artifact detector, evaluated on the original 24-slide partition and against a supervised baseline. Results: The method's central mechanism reproduces: the auxiliary contrastive term improves pooled F1 from 0.673 to 0.688 and replicates under a second seed (+0.0156, p = 0.031; +0.0190, p = 0.005), although it acts on pen marking rather than the artifact types cited to motivate it. Its comparative claims do not: differences between design variants, and against the supervised baseline, fall inside the uncertainty of the evaluation. Four of 24 slides carry 70% of scored annotated pixels, giving an effective sample size of 6.2, and the inherited partition sits at the 7th percentile. An unreported tissue-restriction step excludes 41.4% of out-of-focus annotation against 2.6% of air bubble; such a gate is confounded with blur by construction. Conclusions: Small-cohort benchmarks support far weaker conclusions than current reporting implies. The four checks are cheap enough to accompany any evaluation on such a resource and separate reproducible effects from differences the evaluation cannot resolve.

病理图像评估协议可复现性小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。