arXiv:2609.02663cs.CV2026-09

揭示医学图文分割中文本依赖程度差异及作用机制

Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

论文配图:Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling
图 1 · 摘自论文原文
  • 提出证据解耦解码器,分离图像与文本驱动的分割证据
  • 不同数据集上文本移除导致性能骤降或轻微波动
  • 文本主要通过全局语义调节影响结果,非局部定位

预训练视觉语言模型(VLMs)在结合临床文本后,在医学图像分割任务中表现优异。然而,文本信息对像素级预测的实际贡献仍不明确。本文系统研究了文本在多模态医学图像分割中的作用。首先分析多种常见融合策略,发现分割性能对融合模块选择不敏感。为进一步理解模态交互,提出基于证据深度学习和深度监督的证据解耦解码器(EDD),可分解解码过程中图像证据与文本调制证据,同时保持良好分割性能。实验表明,文本扰动敏感性在不同数据集间差异显著:在BUSI和BTMRI上移除文本导致性能灾难性下降,表明模型高度依赖文本;而在ISIC和Kvasir-SEG上影响较小。进一步发现,文本主要通过全局语义调节而非独立空间定位影响预测,且驱动文本敏感性的具体语义成分因数据集而异。该研究深化了对多模态医学分割中模态交互的理解,并为未来模型设计提供实践指导。

原文摘要 · Abstract (English)

Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.

医学图像分割图文模型模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。