arXiv:2509.25863cs.CV2025-09NeurIPS被引 1

通过多尺度属性提示学习,提升病理切片少样本分类精度

MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification

  • 用大语言模型生成实体与切片级提示,捕捉多尺度病理特征
  • 在三个癌症队列上实现优于基线的方法,提升少样本诊断效果
  • 适合需要少标注的医学图像分析研究者使用

提示学习已成为适应预训练视觉-语言模型(VLMs)于少样本全切片图像(WSI)分类的有前景范式,通过对齐视觉特征与文本表示,降低标注成本并增强模型泛化能力。然而,现有方法通常依赖切片级提示,难以捕捉组织学实体(如细胞核、腺体)的亚型特异性表型差异,而这些差异对癌症诊断至关重要。为此,我们提出多尺度属性增强提示学习(MAPLE),一种分层框架,同时整合多尺度视觉语义,并在实体和切片两个层面进行预测。首先,利用大语言模型(LLMs)生成实体级提示以识别多尺度组织学实体及其表型属性,以及切片级提示以捕捉全局视觉描述。随后,设计实体引导的跨注意力模块生成实体级特征,并与对应亚型特异性属性对齐,实现细粒度实体级预测。为丰富实体表示,进一步开发跨尺度实体图学习模块,通过捕捉不同尺度内及跨尺度间的语义关联来更新表示。最终,将优化后的表示聚合为切片级表示,并与相应提示对齐以完成切片级预测。结合实体级与切片级输出得到最终结果。在三个癌症队列上的实验验证了该方法在解决少样本病理诊断任务中的有效性。

原文摘要 · Abstract (English)

Prompt learning has emerged as a promising paradigm for adapting pre-trained vision-language models (VLMs) to few-shot whole slide image (WSI) classification by aligning visual features with textual representations, thereby reducing annotation cost and enhancing model generalization. Nevertheless, existing methods typically rely on slide-level prompts and fail to capture the subtype-specific phenotypic variations of histological entities (\emph{e.g.,} nuclei, glands) that are critical for cancer diagnosis. To address this gap, we propose Multi-scale Attribute-enhanced Prompt Learning (\textbf{MAPLE}), a hierarchical framework for few-shot WSI classification that jointly integrates multi-scale visual semantics and performs prediction at both the entity and slide levels. Specifically, we first leverage large language models (LLMs) to generate entity-level prompts that can help identify multi-scale histological entities and their phenotypic attributes, as well as slide-level prompts to capture global visual descriptions. Then, an entity-guided cross-attention module is proposed to generate entity-level features, followed by aligning with their corresponding subtype-specific attributes for fine-grained entity-level prediction. To enrich entity representations, we further develop a cross-scale entity graph learning module that can update these representations by capturing their semantic correlations within and across scales. The refined representations are then aggregated into a slide-level representation and aligned with the corresponding prompts for slide-level prediction. Finally, we combine both entity-level and slide-level outputs to produce the final prediction results. Results on three cancer cohorts confirm the effectiveness of our approach in addressing few-shot pathology diagnosis tasks.

少样本学习病理图像提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。