arXiv:2604.09169cs.CV2026-04中稿 · CVPR被引 2

用文本原型对齐提升病理图像分割,仅需10%标注数据就显著更准。

UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation

论文配图:UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation
图 1 · 摘自论文原文
  • 引入文本与原型双对齐分支,让模型学得更清晰的类别结构
  • 在GlaS和CRAG上仅用10%标注数据,Dice提升达8.6%(最大)
  • 适合标注稀缺的病理图像分割任务,尤其适合资源有限的研究者

计算病理学中的半监督语义分割面临标注稀疏和伪标签不可靠的挑战。我们提出UniSemAlign,一种基于病理预训练Transformer编码器的双模态语义对齐框架,通过显式注入类别级结构来增强像素级分割。该框架在共享嵌入空间中引入互补的原型级和文本级对齐分支,提供结构化指导以减少类别歧义并稳定伪标签优化。对齐表示与视觉预测融合,生成更可靠的未标注图像监督信号。框架通过有监督分割、跨视图一致性及跨模态对齐目标端到端训练。在GlaS和CRAG数据集上的大量实验表明,仅使用10%标注数据时,其在GlaS上最多提升2.6%的Dice,在CRAG上最多提升8.6%,20%标注下也表现强劲。

原文摘要 · Abstract (English)

Semi-supervised semantic segmentation in computational pathology remains challenging due to scarce pixel-level annotations and unreliable pseudo-label supervision. We propose UniSemAlign, a dual-modal semantic alignment framework that enhances visual segmentation by injecting explicit class-level structure into pixel-wise learning. Built upon a pathology-pretrained Transformer encoder, UniSemAlign introduces complementary prototype-level and text-level alignment branches in a shared embedding space, providing structured guidance that reduces class ambiguity and stabilizes pseudo-label refinement. The aligned representations are fused with visual predictions to generate more reliable supervision for unlabeled histopathology images. The framework is trained end-to-end with supervised segmentation, cross-view consistency, and cross-modal alignment objectives. Extensive experiments on the GlaS and CRAG datasets demonstrate that UniSemAlign substantially outperforms recent semi-supervised baselines under limited supervision, achieving Dice improvements of up to 2.6% on GlaS and 8.6% on CRAG with only 10% labeled data, and strong improvements at 20% supervision. Code is available at: https://github.com/thailevann/UniSemAlign

病理分割半监督原型对齐文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。