用虚拟染色和沙普利值融合,让病理图像分类更准且可解释。
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
- 通过虚拟染色生成新模态,增强有限病理图像数据。
- 沙普利值筛选关键特征,使多模态融合提升准确率5%-11%。
- 适合需要可解释性的医学影像分析研究者使用。
多模态模型在组织病理学中展现出潜力,但多数基于H&E和基因组数据,设计日益复杂且为黑箱。本文提出新型可解释多模态框架SHAP-CAT,采用基于沙普利值的降维技术实现有效融合。以配对的H&E与IHC图像为输入,利用虚拟染色技术生成新的临床相关模态,增强有限数据。从图像模态中提取轻量级袋级表征,并通过沙普利值计算各维度的贡献度,据此选择每模态中最重要的若干维度进行后期融合。实验表明,引入合成模态的SHAP-CAT显著提升性能,在BCI上准确率提高5%,在IHC4BC-ER上提高8%,在IHC4BC-PR上提高11%。
原文摘要 · Abstract (English)
The multimodal model has demonstrated promise in histopathology. However, most multimodal models are based on H\&E and genomics, adopting increasingly complex yet black-box designs. In our paper, we propose a novel interpretable multimodal framework named SHAP-CAT, which uses a Shapley-value-based dimension reduction technique for effective multimodal fusion. Starting with two paired modalities -- H\&E and IHC images, we employ virtual staining techniques to enhance limited input data by generating a new clinical-related modality. Lightweight bag-level representations are extracted from image modalities and a Shapley-value-based mechanism is used for dimension reduction. For each dimension of the bag-level representation, attribution values are calculated to indicate how changes in the specific dimensions of the input affect the model output. In this way, we select a few top important dimensions of bag-level representation for each image modality to late fusion. Our experimental results demonstrate that the proposed SHAP-CAT framework incorporating synthetic modalities significantly enhances model performance, yielding a 5\% increase in accuracy for the BCI, an 8\% increase for IHC4BC-ER, and an 11\% increase for the IHC4BC-PR dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。