arXiv:2412.02978cs.CV2024-12被引 1

单分支模型融合视觉与语言信息,提升多器官多细胞类型分割精度

Progressive Vision-Language Prompt for Multi-Organ Multi-Class Cell Semantic Segmentation with Single Branch

  • 采用分层特征提取与渐进式提示解码,融合多粒度视觉与文本信息
  • 在PanNuke数据集上达到最优性能,尤其擅长处理形状细微差异的细胞
  • 架构简洁(单分支),适合资源有限场景下的病理图像分析

病理细胞语义分割是计算病理学的基础技术,对癌症诊断和治疗具有重要意义。由于不同器官中存在多种细胞类型,且细胞大小、形状差异细微,多器官、多类别细胞分割极具挑战性。现有方法多采用多分支结构增强特征提取,但导致模型复杂;且过度依赖视觉信息,在多类别分析中受限于复杂的纹理细节。为此,我们提出一种基于单分支的多器官多类别细胞语义分割方法MONCH,通过视觉-语言输入实现跨模态融合。具体而言,设计分层特征提取机制,提供从粗到细的特征表示,涵盖高频、卷积及拓扑特征。受文本与多粒度视觉特征协同启发,引入渐进式提示解码器,自细粒度至粗粒度整合多模态特征,提升上下文建模能力。在具有显著类别不平衡和细微细胞形态差异的PanNuke数据集上,实验表明MONCH优于当前最先进的细胞分割方法与视觉-语言模型。代码与实现将公开。

原文摘要 · Abstract (English)

Pathological cell semantic segmentation is a fundamental technology in computational pathology, essential for applications like cancer diagnosis and effective treatment. Given that multiple cell types exist across various organs, with subtle differences in cell size and shape, multi-organ, multi-class cell segmentation is particularly challenging. Most existing methods employ multi-branch frameworks to enhance feature extraction, but often result in complex architectures. Moreover, reliance on visual information limits performance in multi-class analysis due to intricate textural details. To address these challenges, we propose a Multi-OrgaN multi-Class cell semantic segmentation method with a single brancH (MONCH) that leverages vision-language input. Specifically, we design a hierarchical feature extraction mechanism to provide coarse-to-fine-grained features for segmenting cells of various shapes, including high-frequency, convolutional, and topological features. Inspired by the synergy of textual and multi-grained visual features, we introduce a progressive prompt decoder to harmonize multimodal information, integrating features from fine to coarse granularity for better context capture. Extensive experiments on the PanNuke dataset, which has significant class imbalance and subtle cell size and shape variations, demonstrate that MONCH outperforms state-of-the-art cell segmentation methods and vision-language models. Codes and implementations will be made publicly available.

细胞分割视觉语言单分支病理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。