arXiv:2503.00915cs.CVcs.AI2025-03被引 2

用多模态知识蒸馏提升病理切片分类,解决少数类别样本不足问题。

Multimodal Distillation-Driven Ensemble Learning for Long-Tailed Histopathology Whole Slide Images Analysis

  • 构建多专家集成框架,共享聚合器并约束一致性,学习多样分布。
  • 在Camelyon+-LT和PANDA-LT上准确率超现有方法,少数类识别显著提升。
  • 适合处理标注不均的医学图像分析,尤其对小样本类别有优势。

多实例学习(MIL)在计算病理学中至关重要,支持弱监督下的全切片图像(WSI)分析。然而,WSI分析面临严重的长尾分布问题,导致类别不平衡:部分类别样本稀少,另一些则数量众多,影响分类器对少数类的识别能力。为此,本文提出基于MIL的集成学习方法MDE-MIL,采用具有共享聚合器和一致性约束的专家解码器,以学习多样化分布并缓解类别不平衡的影响。进一步引入多模态知识蒸馏框架,利用在病理-文本对上预训练的文本编码器,引导MIL聚合器捕捉更相关的语义特征。为保证灵活性,使用可学习提示控制蒸馏过程,避免固定提示带来的限制。MDE-MIL通过多个聚焦特定数据分布的专家分支,结合一致性控制与多模态蒸馏,有效提升特征提取能力。在Camelyon+-LT和PANDA-LT数据集上的实验表明,该方法优于当前最先进方法。

原文摘要 · Abstract (English)

Multiple Instance Learning (MIL) plays a significant role in computational pathology, enabling weakly supervised analysis of Whole Slide Image (WSI) datasets. The field of WSI analysis is confronted with a severe long-tailed distribution problem, which significantly impacts the performance of classifiers. Long-tailed distributions lead to class imbalance, where some classes have sparse samples while others are abundant, making it difficult for classifiers to accurately identify minority class samples. To address this issue, we propose an ensemble learning method based on MIL, which employs expert decoders with shared aggregators and consistency constraints to learn diverse distributions and reduce the impact of class imbalance on classifier performance. Moreover, we introduce a multimodal distillation framework that leverages text encoders pre-trained on pathology-text pairs to distill knowledge and guide the MIL aggregator in capturing stronger semantic features relevant to class information. To ensure flexibility, we use learnable prompts to guide the distillation process of the pre-trained text encoder, avoiding limitations imposed by specific prompts. Our method, MDE-MIL, integrates multiple expert branches focusing on specific data distributions to address long-tailed issues. Consistency control ensures generalization across classes. Multimodal distillation enhances feature extraction. Experiments on Camelyon+-LT and PANDA-LT datasets show it outperforms state-of-the-art methods.

病理图像长尾分布知识蒸馏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。