arXiv:2608.14766cs.CVcs.LG2026-08被引 1

医学影像中,传统不确定性估计无法捕捉病灶是否存在这一关键模糊性。

Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation

论文配图:Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation
图 1 · 摘自论文原文
  • 用冻结特征+轻量监督头替代熵聚合,更有效识别病灶存在性模糊
  • 在4种模型、2个数据集上均显著优于基于熵的不确定性方法
  • 揭示了现有方法忽略编码器中已存在的存在性模糊信息

不确定性估计对深度学习在医疗图像分割中的安全临床应用至关重要,其中似然不确定性理论上应捕捉不可消除的数据模糊性。然而,基于熵的度量是否反映临床有意义的模糊性——即病例层面关于病灶是否存在的一致性分歧——仍不明确。与以往关注像素级边界分歧的研究不同,本文系统评估了似然不确定性在捕捉存在性模糊方面的能力。研究涵盖四种架构,使用蒙特卡洛丢弃和深度集成,在LIDC-IDRI和外部验证队列LNDb上进行3D肺结节分割。结果表明,基于熵的不确定性图仅与边界噪声和微小勾画差异相关,对存在性模糊缺乏判别力。相反,基于冻结分割特征训练的轻量级监督模糊性头,在所有架构、指标和数据集上均显著优于所有基于熵聚合的基线方法,且表现可媲美或超过需显式建模模糊性的方法(如Probabilistic U-Net、Annotator-Confusion 3D-UNet)。定性特征空间分析显示,存在性模糊已在像素级训练网络的冻结编码器特征中编码,却被分割输出及其熵聚合过程丢弃。研究揭示了似然不确定性理论承诺与其实际行为间的根本不匹配,提示从业者不应将基于熵的不确定性作为安全关键应用中临床模糊性的代理。

原文摘要 · Abstract (English)

Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. case-level disagreement about whether a pathology is present at all, remains poorly understood. Contrary to most prior work, which focused on pixel-wise boundary disagreement, we systematically evaluate how well aleatoric uncertainty captures presence ambiguity. Our evaluation spans 3D lung nodule segmentation across four architectures with Monte Carlo dropout and deep ensembles, on LIDC-IDRI and an external validation cohort (LNDb). We find that entropy-based uncertainty maps align with boundary noise and minor drawing variation but carry insufficient discriminative signal for presence ambiguity. In contrast, a lightweight supervised ambiguity head trained on frozen segmentation features substantially outperforms all entropy-aggregation-based baselines across architectures, metrics, and both cohorts, and matches or exceeds methods that explicitly model ambiguity under disagreement supervision (Probabilistic U-Net, Annotator-Confusion 3D-UNet). A qualitative feature-space analysis shows that presence ambiguity is already encoded in the frozen encoder features of pixel-wise-trained networks, only to be discarded by the segmentation output and its entropy aggregation. Our findings expose a fundamental mismatch between the theoretical promise of aleatoric uncertainty and its practical behavior, and suggest that practitioners should not rely on entropy-based uncertainty as a proxy for clinical ambiguity in safety-critical applications.

医学图像不确定性估计肺结节深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。