arXiv:2604.15950cs.LG2026-04中稿 · publication at MID…

让医学影像分割结果直接反映多位医生的共识程度。

TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation

论文配图:TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation
图 1 · 摘自论文原文
  • 用多位医生标注的平均比例校准模型输出概率
  • 在CURVAS-PDACVI数据集上显著提升校准度
  • 适合需要解释性医疗决策的场景

增强CT图像上胰腺导管腺癌(PDAC)分割具有内在模糊性:专家间标注差异反映的是真实不确定性而非标注噪声。标准深度学习方法假设存在单一真实标签,导致概率输出校准不佳且难以解释。本文提出TwinTrack框架,通过后处理校准集成分割概率,使其匹配经验性的人类平均响应(MHR)——即标记某个体素为肿瘤的专家占比。校准后的概率可直接解读为预期标注者中赋予肿瘤标签的比例,显式建模了专家间分歧。该后处理校准方法简单,仅需少量多标注者校准数据集,且在MICCAI 2025 CURVAS-PDACVI多标注者基准测试中持续优于标准方法。

原文摘要 · Abstract (English)

Pancreatic ductal adenocarcinoma (PDAC) segmentation on contrast-enhanced CT is inherently ambiguous: inter-rater disagreement among experts reflects genuine uncertainty rather than annotation noise. Standard deep learning approaches assume a single ground truth, producing probabilistic outputs that can be poorly calibrated and difficult to interpret under such ambiguity. We present TwinTrack, a framework that addresses this gap through post-hoc calibration of ensemble segmentation probabilities to the empirical mean human response (MHR) -the fraction of expert annotators labeling a voxel as tumor. Calibrated probabilities are thus directly interpretable as the expected proportion of annotators assigning the tumor label, explicitly modeling inter-rater disagreement. The proposed post-hoc calibration procedure is simple and requires only a small multi-rater calibration set. It consistently improves calibration metrics over standard approaches when evaluated on the MICCAI 2025 CURVAS-PDACVI multi-rater benchmark.

医学分割不确定性建模模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。