用视觉语言模型提升无配对病理切片虚拟染色的准确性与真实性
VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
- 引入病理视觉语言模型作为辅助,结合可学习提示与组织/染色概念锚点
- 在多个公开数据集上实现更逼真的虚拟染色图像,提升肾小球检测精度
- 适合关注病理图像生成、医学视觉推理与跨模态融合的研究者
在病理学中,组织切片通常采用H&E染色或特殊染色(如MAS、PAS、PASM等)以清晰显示特定组织结构。深度学习的快速发展为生成虚拟染色图像提供了有效方案,显著降低了传统化学染色的时间和人力成本。然而,如何分离组织切片的基础视觉特征与染色剂引起的视觉差异成为新挑战。此外,现有虚拟染色方法常忽略关键病理知识与染色物理特性,仅实现风格层面的迁移。为此,我们首次在虚拟染色任务中引入病理视觉语言大模型(VLM)作为辅助工具,整合对比可学习提示、组织切片基础概念锚点及染色特异性概念锚点,充分挖掘病理VLM中的丰富知识,用于描述、构建并引导虚拟染色方向。同时,我们提出一种基于VLM约束的数据增强方法,利用其强大的图像理解能力,进一步融合图像风格与结构信息,在高精度病理诊断中表现优异。在多个公开的多域无配对染色数据集上的大量评估表明,该方法能生成高度逼真的虚拟染色图像,并显著提升下游任务(如肾小球检测与分割)的性能。代码已开源:https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR。
原文摘要 · Abstract (English)
In histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc.) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution for generating virtually stained images, significantly reducing the time and labor costs associated with traditional histochemical staining. However, a new challenge arises in separating the fundamental visual characteristics of tissue sections from the visual differences induced by staining agents. Additionally, virtual staining often overlooks essential pathological knowledge and the physical properties of staining, resulting in only style-level transfer. To address these issues, we introduce, for the first time in virtual staining tasks, a pathological vision-language large model (VLM) as an auxiliary tool. We integrate contrastive learnable prompts, foundational concept anchors for tissue sections, and staining-specific concept anchors to leverage the extensive knowledge of the pathological VLM. This approach is designed to describe, frame, and enhance the direction of virtual staining. Furthermore, we have developed a data augmentation method based on the constraints of the VLM. This method utilizes the VLM's powerful image interpretation capabilities to further integrate image style and structural information, proving beneficial in high-precision pathological diagnostics. Extensive evaluations on publicly available multi-domain unpaired staining datasets demonstrate that our method can generate highly realistic images and enhance the accuracy of downstream tasks, such as glomerular detection and segmentation. Our code is available at: https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。