用图文原型联合学习,提升病理图像弱监督分割精度
DualProtoSeg: Simple and Efficient Design with Text- and Image-Guided Prototype Learning for Weakly Supervised Histopathology Image Segmentation
- 引入文本与图像双模态原型库,融合语义与外观特征
- 在BCSS-WSSS数据集上超越现有最佳方法,分割准确率显著提升
- 适合关注医学图像分割、少样本学习的研究者
组织学图像的弱监督语义分割旨在通过图像级标签降低标注成本,但仍受限于类别间同质性、类内异质性以及基于CAM的监督导致的区域收缩问题。本文提出一种简单高效的原型驱动框架,利用视觉-语言对齐提升弱监督下的区域发现能力。方法结合CoOp风格的可学习提示调优生成文本原型,并与可学习图像原型融合,构建双模态原型库,同时捕捉语义与外观线索。为缓解ViT表示中的过平滑问题,引入多尺度金字塔模块,增强空间精度与定位质量。在BCSS-WSSS基准上的实验表明,该方法优于现有最先进方法;详细分析验证了文本描述多样性、上下文长度及图文原型互补性的优势。结果表明,联合利用文本语义与视觉原型学习在数字病理弱监督分割中具有显著有效性。
原文摘要 · Abstract (English)
Weakly supervised semantic segmentation (WSSS) in histopathology seeks to reduce annotation cost by learning from image-level labels, yet it remains limited by inter-class homogeneity, intra-class heterogeneity, and the region-shrinkage effect of CAM-based supervision. We propose a simple and effective prototype-driven framework that leverages vision-language alignment to improve region discovery under weak supervision. Our method integrates CoOp-style learnable prompt tuning to generate text-based prototypes and combines them with learnable image prototypes, forming a dual-modal prototype bank that captures both semantic and appearance cues. To address oversmoothing in ViT representations, we incorporate a multi-scale pyramid module that enhances spatial precision and improves localization quality. Experiments on the BCSS-WSSS benchmark show that our approach surpasses existing state-of-the-art methods, and detailed analyses demonstrate the benefits of text description diversity, context length, and the complementary behavior of text and image prototypes. These results highlight the effectiveness of jointly leveraging textual semantics and visual prototype learning for WSSS in digital pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。