用文本提示提升乳腺超声分割精度,尤其改善小病灶边界模糊问题。
XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
- 双路径结构:全局语义+局部边界建模,结合临床文本提示。
- 在BLU数据集上达到0.8765的Dice和0.8149的IoU,小病灶表现最佳。
- 自动提取病灶描述文本,无需人工标注,适合临床部署。
精确的乳腺超声(BUS)分割有助于可靠测量、定量分析和下游分类,但对小或低对比度病灶、边缘模糊及斑点噪声仍具挑战。文本提示可引入临床上下文,但直接使用弱定位的图文信号(如CAM/CLIP生成)常导致粗略、块状响应,边界模糊,除非额外机制恢复精细轮廓。本文提出XBusNet,一种双提示、双分支多模态模型,融合图像特征与临床相关文本。全局路径基于CLIP视觉变换器,编码整体图像语义,受病灶大小和位置条件约束;局部路径采用U-Net,强调精确边界,并由描述形状、边缘和BI-RADS术语的提示调制。提示通过结构化元数据自动组装,无需人工点击。在乳腺病灶超声(BLU)数据集上进行五折交叉验证。主要指标为Dice和交并比(IoU);还开展分尺寸分析与消融实验,评估全局路径、局部路径及文本调制的作用。结果表明,XBusNet在BLU上达到最新性能,平均Dice为0.8765,IoU为0.8149,优于六种强基线模型。小病灶获益最大,漏检减少,误激活降低。消融实验显示全局上下文、局部边界建模与文本调制具有互补贡献。结论:双提示、双分支多模态设计融合全局语义与局部精度,实现精准的乳腺超声分割,增强对小而低对比病灶的鲁棒性。
原文摘要 · Abstract (English)
Background: Precise breast ultrasound (BUS) segmentation supports reliable measurement, quantitative analysis, and downstream classification, yet remains difficult for small or low-contrast lesions with fuzzy margins and speckle noise. Text prompts can add clinical context, but directly applying weakly localized text-image cues (e.g., CAM/CLIP-derived signals) tends to produce coarse, blob-like responses that smear boundaries unless additional mechanisms recover fine edges. Methods: We propose XBusNet, a novel dual-prompt, dual-branch multimodal model that combines image features with clinically grounded text. A global pathway based on a CLIP Vision Transformer encodes whole-image semantics conditioned on lesion size and location, while a local U-Net pathway emphasizes precise boundaries and is modulated by prompts that describe shape, margin, and Breast Imaging Reporting and Data System (BI-RADS) terms. Prompts are assembled automatically from structured metadata, requiring no manual clicks. We evaluate on the Breast Lesions USG (BLU) dataset using five-fold cross-validation. Primary metrics are Dice and Intersection over Union (IoU); we also conduct size-stratified analyses and ablations to assess the roles of the global and local paths and the text-driven modulation. Results: XBusNet achieves state-of-the-art performance on BLU, with mean Dice of 0.8765 and IoU of 0.8149, outperforming six strong baselines. Small lesions show the largest gains, with fewer missed regions and fewer spurious activations. Ablation studies show complementary contributions of global context, local boundary modeling, and prompt-based modulation. Conclusions: A dual-prompt, dual-branch multimodal design that merges global semantics with local precision yields accurate BUS segmentation masks and improves robustness for small, low-contrast lesions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。