用视觉变压器+扩散模型,精准分割乳腺超声中的病灶。
A Dual-Mode ViT-Conditioned Diffusion Framework with an Adaptive Conditioning Bridge for Breast Cancer Segmentation
- 结合ViT与扩散模型,通过自适应桥接融合多尺度特征。
- 在三个公开数据集上达到0.97、0.96、0.90的骰子分数。
- 适合需要高精度与解剖合理性分割的医学影像研究者。
在乳腺超声图像中,精确分割病灶对早期诊断至关重要;然而,低对比度、斑点噪声和边界模糊使其困难。尽管深度学习模型展现出潜力,但标准卷积架构常因缺乏全局上下文而产生解剖不一致的分割结果。为此,我们提出一种灵活的条件去噪扩散模型,结合增强型基于UNet的生成解码器与视觉变压器(ViT)编码器以提取全局特征。引入三项核心创新:1)自适应条件桥(ACB),实现语义特征的高效多尺度融合;2)新型拓扑去噪一致性(TDC)损失,通过惩罚去噪过程中的结构不一致来正则化训练;3)双头架构,利用去噪目标作为强正则化器,使轻量级辅助头可在小数据集上实现快速准确推理,同时保留噪声预测头。该框架在公开乳腺超声数据集上达到新最优性能,于BUSI、BrEaST和BUS-UCLM上分别取得0.96、0.90和0.97的骰子分数。全面的消融实验验证了各组件对性能与解剖合理性的重要性。
原文摘要 · Abstract (English)
In breast ultrasound images, precise lesion segmentation is essential for early diagnosis; however, low contrast, speckle noise, and unclear boundaries make this difficult. Even though deep learning models have demonstrated potential, standard convolutional architectures frequently fall short in capturing enough global context, resulting in segmentations that are anatomically inconsistent. To overcome these drawbacks, we suggest a flexible, conditional Denoising Diffusion Model that combines an enhanced UNet-based generative decoder with a Vision Transformer (ViT) encoder for global feature extraction. We introduce three primary innovations: 1) an Adaptive Conditioning Bridge (ACB) for efficient, multi-scale fusion of semantic features; 2) a novel Topological Denoising Consistency (TDC) loss component that regularizes training by penalizing structural inconsistencies during denoising; and 3) a dual-head architecture that leverages the denoising objective as a powerful regularizer, enabling a lightweight auxiliary head to perform rapid and accurate inference on smaller datasets and a noise prediction head. Our framework establishes a new state-of-the-art on public breast ultrasound datasets, achieving Dice scores of 0.96 on BUSI, 0.90 on BrEaST and 0.97 on BUS-UCLM. Comprehensive ablation studies empirically validate that the model components are critical for achieving these results and for producing segmentations that are not only accurate but also anatomically plausible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。