通过图文细粒度对齐,提升病理图像零样本肿瘤分型准确率
Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- 用局部特征增强模块捕捉病理切片空间关系
- 大模型生成病灶级语义原型,提升类别区分度
- 在EBRAINS和TCGA数据集上达最优零样本表现
从组织病理全切片图像中进行脑肿瘤亚型的细粒度分类极具挑战,源于形态学差异微小且标注数据稀缺。尽管视觉-语言模型已实现有前景的零样本分类,但其对细微病理特征的捕捉能力仍有限,导致亚型判别效果不佳。为此,我们提出细粒度切片-文本对齐网络(FG-PAN),一种专为数字病理设计的零样本框架。FG-PAN包含两个关键模块:(1) 局部特征精炼模块,通过建模代表性切片间的空间关系,增强切片级视觉特征;(2) 细粒度文本描述生成模块,利用大语言模型生成具有病理感知、类别特异的语义原型。通过将精炼后的视觉特征与大模型生成的细粒度描述对齐,FG-PAN有效提升了视觉与语义空间中的类别可分性。在多个公开病理数据集(包括EBRAINS和TCGA)上的大量实验表明,FG-PAN在零样本脑肿瘤亚型分类中达到当前最优性能并具备强泛化能力。
原文摘要 · Abstract (English)
The fine-grained classification of brain tumor subtypes from histopathological whole slide images is highly challenging due to subtle morphological variations and the scarcity of annotated data. Although vision-language models have enabled promising zero-shot classification, their ability to capture fine-grained pathological features remains limited, resulting in suboptimal subtype discrimination. To address these challenges, we propose the Fine-Grained Patch Alignment Network (FG-PAN), a novel zero-shot framework tailored for digital pathology. FG-PAN consists of two key modules: (1) a local feature refinement module that enhances patch-level visual features by modeling spatial relationships among representative patches, and (2) a fine-grained text description generation module that leverages large language models to produce pathology-aware, class-specific semantic prototypes. By aligning refined visual features with LLM-generated fine-grained descriptions, FG-PAN effectively increases class separability in both visual and semantic spaces. Extensive experiments on multiple public pathology datasets, including EBRAINS and TCGA, demonstrate that FG-PAN achieves state-of-the-art performance and robust generalization in zero-shot brain tumor subtype classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。