arXiv:2608.28716eess.IVcs.CV2026-08中稿 · as a CaPTion works…

研究影像分割误差对癌症评分决策的影响,发现多数情况不影响判断,边缘案例仍需专家审核。

Evaluating the Effects of Inter-Observer and Model Variability on Radiological Peritoneal Cancer Index Assessment

论文配图:Evaluating the Effects of Inter-Observer and Model Variability on Radiological Peritoneal Cancer Index Assessment
图 1 · 摘自论文原文
  • 用真实医生标注作参考,对比模型与人对腹膜癌区域的分割差异。
  • 模型在多数区域表现接近医生,但小肠区和特定区域偏差较大。
  • 评分差异小,仅在临界值附近易导致诊断变化,适合临床部署参考。

深度学习分割模型通常以Dice、HD95和ASD等几何指标评估,但这些指标提升是否带来临床决策改善尚不明确。本文通过增强CT上的放射科腹膜癌指数(rPCI)分割,利用共识定义的13个解剖三维区域及临床常用的PCI 20阈值,评估指标到决策的差距。四位专家在十例腹部CT上进行标注,量化观察者间变异;并以公开的nnU-Net为基础的rPCI分割模型为基准,计算各区域的Dice、HD95和ASD。为连接几何差异与临床影响,构建概率性腹膜转移模拟,在多数投票的rPCI图上传播边界变异,生成衍生的(r)PCI评分及在PCI 20阈值下的分类结果。结果显示,观察者间一致性高(平均Dice为0.87),模型在多数区域匹配人类表现,但在区域4、8以及小肠区域(9–12)偏差较大。模拟显示,评分差异普遍较小(均值ΔrPCI ≈ 0.3–0.6),决策翻转主要发生在参考评分接近20时。表明rPCI评分对典型分割变异具有鲁棒性,而临界病例仍是专家审核的关键场景。

原文摘要 · Abstract (English)

Deep learning segmentation models are often evaluated using geometric metrics such as Dice, HD95, and ASD, yet it remains unclear to what extent improvements in these metrics translate into clinically meaningful changes in downstream decision-making. The metric-to-decision gap is examined using radiological Peritoneal Cancer Index (rPCI) region segmentation on contrast-enhanced CT, where a consensus definition provides anatomically grounded 3D regions and the clinically used PCI 20 threshold enables decision-level evaluation. Inter-observer variability is quantified across four experts on ten abdominal CT scans, and a published nnU-Net based rPCI segmentation model is benchmarked against this human reference using Dice, HD95, and ASD across all 13 regions. To relate geometric differences to clinical impact, a probabilistic peritoneal metastasis simulation is implemented on majority-vote rPCI maps, propagating region-boundary variability into variability of derived (r)PCI scores and classification at the PCI 20 cutoff. Observers showed high agreement (mean Dice $0.87$), while the model matched human performance in most regions but deviated more in regions 4, 8, and the small-bowel regions (9-12). Across simulations, score differences were typically small (mean $Δ$rPCI $\approx 0.3$-$0.6$) for both observers and the model, and decision flips occurred predominantly when the reference score was near 20. These results suggest that rPCI-derived scoring is generally robust to typical segmentation variability, while highlighting borderline cases as the main setting where expert review remains essential.

医学影像分割评估临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。