arXiv:2604.18713cs.CV2026-04中稿 · EMBC 2026

用文本引导精准分割前列腺病变,提升多模态MRI分析精度。

Align then Refine: Text-Guided 3D Prostate Lesion Segmentation

论文配图:Align then Refine: Text-Guided 3D Prostate Lesion Segmentation
图 1 · 摘自论文原文
  • 先对齐文本与图像特征,再通过注意力机制精细修正边界。
  • 在PI-CAI数据集上达到新最好结果,平均Dice达0.827。
  • 适合医学影像分析、多模态融合与细粒度语义分割研究者。

从双参数MRI(bp-MRI)自动分割3D前列腺病变对可靠算法分析至关重要,但高精度仍具挑战。体积化方法需融合多模态信息并保证解剖一致性,而现有模型难以可靠整合跨模态信息。尽管视觉语言模型(VLMs)正在取代传统架构,仍缺乏用于有效局部引导的细粒度病变级语义。为此,我们提出一种新型多编码器U-Net架构,包含三项创新:(1) 对齐损失增强前景文本-图像相似性,注入病变语义;(2) 热图损失校准相似性图并抑制伪背景激活;(3) 高置信度区域的置信度门控多头交叉注意力精修器,实现局部边界优化。分阶段训练策略稳定了各组件优化。本方法在PI-CAI数据集上持续超越先前方法,通过增强多模态融合与局部文本引导,确立新基准。代码已开源:https://github.com/NUBagciLab/Prostate-Lesion-Segmentation。

原文摘要 · Abstract (English)

Automated 3D segmentation of prostate lesions from biparametric MRI (bp-MRI) is essential for reliable algorithmic analysis, but achieving high precision remains challenging. Volumetric methods must combine multiple modalities while ensuring anatomical consistency, but current models struggle to integrate cross-modal information reliably. While vision-language models (VLMs) are replacing the currently used architectural designs, they still lack the fine-grained, lesion-level semantics required for effective localized guidance. To address these limitations, we propose a new multi-encoder U-Net architecture incorporating three key innovations: (1) an alignment loss that enhances foreground text-image similarity to inject lesion semantics; (2) a heatmap loss that calibrates the similarity map and suppresses spurious background activations; and (3) a final-stage, confidence-gated multi-head cross-attention refiner that performs localized boundary edits in high-confidence regions. A phase-scheduled training regime stabilizes the optimization of these components. Our method consistently outperforms prior approaches, establishing a new state-of-the-art on the PI-CAI dataset through enhanced multi-modal fusion and localized text guidance. Our code is available at https://github.com/NUBagciLab/Prostate-Lesion-Segmentation.

前列腺分割多模态融合文本引导医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。