arXiv:2512.10316cs.CV2025-12

用视觉语言模型和分割网络融合生成病理图像伪掩码,提升弱监督分割精度。

ConStruct: Structural Distillation of Foundation Models for Prototype-Based Weakly Supervised Histopathology Segmentation

  • 融合CONCH语义与SegFormer结构,生成兼具语义和空间一致性的原型
  • 通过文本引导初始化和结构蒸馏,提升伪掩码完整性和边界清晰度
  • 无需像素级标注,适合缺乏精细标注的病理图像分割任务

病理图像弱监督语义分割依赖分类主干网络,但这些模型常仅定位最具判别性的区域,难以捕捉组织结构的完整空间范围。视觉语言模型如CONCH具备丰富的语义对齐和形态感知表征,而现代分割主干如SegFormer能保留细粒度空间线索。然而,在弱监督且无密集标注条件下,整合二者优势仍具挑战。本文提出一种原型学习框架,融合CONCH的形态感知表征、SegFormer的多尺度结构线索及文本引导的语义对齐,生成同时具备语义判别力和空间一致性的原型。为有效利用异构来源,引入文本引导的原型初始化,结合病理描述生成更完整、语义准确的伪掩码;设计结构蒸馏机制,将SegFormer的空间知识迁移至原型学习过程,以保留细粒度形态模式和局部组织边界。实验在BCSS-WSSS数据集上表明,该方法在不依赖像素级标注的情况下,显著提升定位完整性与跨组织类型的语义一致性,优于现有弱监督分割方法,且通过冻结基础模型主干与轻量适配器实现高效计算。

原文摘要 · Abstract (English)

Weakly supervised semantic segmentation (WSSS) in histopathology relies heavily on classification backbones, yet these models often localize only the most discriminative regions and struggle to capture the full spatial extent of tissue structures. Vision-language models such as CONCH offer rich semantic alignment and morphology-aware representations, while modern segmentation backbones like SegFormer preserve fine-grained spatial cues. However, combining these complementary strengths remains challenging, especially under weak supervision and without dense annotations. We propose a prototype learning framework for WSSS in histopathological images that integrates morphology-aware representations from CONCH, multi-scale structural cues from SegFormer, and text-guided semantic alignment to produce prototypes that are simultaneously semantically discriminative and spatially coherent. To effectively leverage these heterogeneous sources, we introduce text-guided prototype initialization that incorporates pathology descriptions to generate more complete and semantically accurate pseudo-masks. A structural distillation mechanism transfers spatial knowledge from SegFormer to preserve fine-grained morphological patterns and local tissue boundaries during prototype learning. Our approach produces high-quality pseudo masks without pixel-level annotations, improves localization completeness, and enhances semantic consistency across tissue types. Experiments on BCSS-WSSS datasets demonstrate that our prototype learning framework outperforms existing WSSS methods while remaining computationally efficient through frozen foundation model backbones and lightweight trainable adapters.

弱监督分割病理图像原型学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。