用语义聚合提升医学图像分割的跨域泛化能力
Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
- 提出EM聚类机制动态凝聚视觉特征,减少模态间偏差
- 文本引导解码器利用通用文本知识对齐医学图像细节
- 在心脏与眼底数据集上均超越现有最优方法
多模态模型在自然图像分割中表现优异,但在医学领域常表现不佳。我们发现其主要瓶颈在于文本提示与细粒度医学视觉特征之间的语义鸿沟,以及由此导致的特征分散。为此,我们从语义聚合角度出发,提出期望-最大化(EM)聚类机制与文本引导像素解码器。前者通过动态将特征聚类为紧凑语义中心,增强跨模态对应关系;后者利用领域无关的文本知识,有效引导深层视觉表征以弥合语义差距。二者协同显著提升了模型泛化能力。在公开的心脏与眼底数据集上的大量实验表明,该方法在多个域泛化基准下持续优于现有最先进方法。
原文摘要 · Abstract (English)
Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal fusion, primarily the significant semantic gap between abstract textual prompts and fine-grained medical visual features, as well as the resulting feature dispersion. To address these issues, we revisit the problem from the perspective of semantic aggregation. Specifically, we propose an Expectation-Maximization (EM) Aggregation mechanism and a Text-Guided Pixel Decoder. The former mitigates feature dispersion by dynamically clustering features into compact semantic centers to enhance cross-modal correspondence. The latter is designed to bridge the semantic gap by leveraging domain-invariant textual knowledge to effectively guide deep visual representations. The synergy between these two mechanisms significantly improves the model's generalization ability. Extensive experiments on public cardiac and fundus datasets demonstrate that our method consistently outperforms existing SOTA approaches across multiple domain generalization benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。