arXiv:2511.07941cs.CVcs.AI2025-11AAAI被引 2

用多模态原型和双向融合提升病理切片少样本分类效果

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

  • 构建任务特异的病理实体文本与视觉原型,实现双向跨模态交互
  • 在三个癌症数据集上少样本分类准确率显著优于基线方法
  • 适合需要可解释性病理分析的医学AI研究者参考

尽管大型语言模型(LLMs)在计算病理学中展现出潜力,但海量像素的全切片图像(WSIs)带来巨大计算负担,需依赖多实例学习(MIL)进行有效建模。典型病理任务仅提供整体袋级标签,而由LLM生成的实例级描述常因缺乏细粒度医学知识存在偏差。为此,我们提出构建任务特定的病理实体原型对学习泛化特征与提升模型可解释性至关重要。现有视觉-语言MIL方法多采用单向引导,限制了跨模态协同。本文提出一种基于多模态原型的MIL方法,通过均衡信息压缩机制促进双向交互。具体地,利用冻结的LLM生成任务特定的病理实体描述,并将其作为文本原型;同时,视觉分支学习实例级原型以减少对冗余数据的依赖。在融合阶段,采用基于相似性度量的立体最优传输(SOT)算法,在高维空间实现更广义的语义对齐。我们在三个不同癌症数据集上进行少样本分类与可解释性实验,结果表明所提方法具有优越的泛化能力。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) are emerging as a promising direction in computational pathology, the substantial computational cost of giga-pixel Whole Slide Images (WSIs) necessitates the use of Multi-Instance Learning (MIL) to enable effective modeling. A key challenge is that pathological tasks typically provide only bag-level labels, while instance-level descriptions generated by LLMs often suffer from bias due to a lack of fine-grained medical knowledge. To address this, we propose that constructing task-specific pathological entity prototypes is crucial for learning generalizable features and enhancing model interpretability. Furthermore, existing vision-language MIL methods often employ unidirectional guidance, limiting cross-modal synergy. In this paper, we introduce a novel approach, Multimodal Prototype-based Multi-Instance Learning, that promotes bidirectional interaction through a balanced information compression scheme. Specifically, we leverage a frozen LLM to generate task-specific pathological entity descriptions, which are learned as text prototypes. Concurrently, the vision branch learns instance-level prototypes to mitigate the model's reliance on redundant data. For the fusion stage, we employ the Stereoscopic Optimal Transport (SOT) algorithm, which is based on a similarity metric, thereby facilitating broader semantic alignment in a higher-dimensional space. We conduct few-shot classification and explainability experiments on three distinct cancer datasets, and the results demonstrate the superior generalization capabilities of our proposed method.

医学影像少样本学习多模态融合可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。