arXiv:2604.10963cs.AI2026-04被引 1

用视觉大模型量化医学影像分割中的固有不确定性,提升模型鲁棒性。

Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

论文配图:Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models
图 1 · 摘自论文原文
  • 通过解码特征的奇异值能量衡量每类样本的语义感知尺度。
  • 在5个数据集上显著提升多器官与肿瘤分割性能,跨架构通用性强。
  • 适合关注医学图像噪声与标注不一致问题的研究者。

医学图像分割支持临床工作流,精准勾画解剖结构与病灶。然而,医学影像数据集存在采集噪声和标注模糊,导致普遍的数据不确定性,严重削弱模型鲁棒性。现有研究多集中于模型结构改进和预测可靠性估计,对内在数据不确定性的系统探索仍不足。为此,本文利用视觉基础模型的通用表征能力,估计固有数据不确定性。具体而言,分析模型解码表示的特征多样性,通过奇异值能量量化每类的语义感知尺度,从而衡量样本难度与后验不确定性。在此基础上,设计两种不确定性驱动的应用策略:(1) 基于不确定性感知的数据过滤机制,剔除潜在噪声样本以提升学习质量;(2) 动态不确定性感知优化策略,依据语义感知尺度自适应调整类别损失权重,并结合标签去噪机制增强训练稳定性。在涵盖CT与MRI模态、涉及多器官与肿瘤分割任务的五个公开数据集上的实验表明,该方法在多种主流网络架构上均实现显著且稳健的性能提升,揭示了后验不确定性在医学图像理解与分割中的广泛应用潜力。

原文摘要 · Abstract (English)

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing pervasive data uncertainty that substantially undermines model robustness. Existing research focuses primarily on model architectural improvements and predictive reliability estimation, while systematic exploration of the intrinsic data uncertainty remains insufficient. To address this gap, this work proposes leveraging the universal representation capabilities of visual foundation models to estimate inherent data uncertainty. Specifically, we analyze the feature diversity of the model's decoded representations and quantify their singular value energy to define the semantic perception scale for each class, thereby measuring sample difficulty and aleatoric uncertainty. Based on this foundation, we design two uncertainty-driven application strategies: (1) the aleatoric uncertainty-aware data filtering mechanism to eliminate potentially noisy samples and enhance model learning quality; (2) the dynamic uncertainty-aware optimization strategy that adaptively adjusts class-specific loss weights during training based on the semantic perception scale, combined with a label denoising mechanism to improve training stability. Experimental results on five public datasets encompassing CT and MRI modalities and involving multi-organ and tumor segmentation tasks demonstrate that our method achieves significant and robust performance improvements across various mainstream network architectures, revealing the broad application potential of aleatoric uncertainty in medical image understanding and segmentation tasks.

医学影像不确定性视觉大模型分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。