arXiv:2503.01904cs.CVcs.AI2025-03被引 2

提出可量化各模态贡献度的新方法,揭示医疗多模态模型的偏倚问题。

What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning

  • 用遮蔽法评估每种模态对模型任务的重要性,不依赖具体模型或性能指标。
  • 发现部分模型存在单一模态偏好,某些数据集本身也存在模态不平衡。
  • 提供细粒度的模态重要性量化与可视化,助力临床可解释性提升。

目的:如今,高维多模态数据可被大型深度神经网络轻松处理,多种模态融合方法已发展成熟。鉴于医学中高维多模态患者数据的普遍性,多模态模型的发展具有重要意义。然而,这些模型如何细致处理各数据源信息仍缺乏研究。方法:为此,我们实现了一种基于遮蔽的模态贡献度测量方法,该方法对模型和性能均无依赖性,能定量评估每个模态在数据集中对模型完成任务的重要性。我们在三个不同的多模态医疗问题上进行了实验验证。结果:我们发现部分网络存在模态偏好,易导致单模态坍缩;同时,某些数据集从源头就存在模态不平衡。此外,我们提供了每种模态的细粒度定量及可视化重要性分析。结论:该度量为多模态模型开发与数据集构建提供了宝贵洞察。通过引入此方法,我们推动了多模态深度学习可解释性研究的发展,有助于多模态AI在临床实践中的整合。代码公开于 https://github.com/ChristianGappGit/MC_MMD。

原文摘要 · Abstract (English)

Purpose High dimensional, multimodal data can nowadays be analyzed by huge deep neural networks with little effort. Several fusion methods for bringing together different modalities have been developed. Given the prevalence of high-dimensional, multimodal patient data in medicine, the development of multimodal models marks a significant advancement. However, how these models process information from individual sources in detail is still underexplored. Methods To this end, we implemented an occlusion-based modality contribution method that is both model- and performance-agnostic. This method quantitatively measures the importance of each modality in the dataset for the model to fulfill its task. We applied our method to three different multimodal medical problems for experimental purposes. Results Herein we found that some networks have modality preferences that tend to unimodal collapses, while some datasets are imbalanced from the ground up. Moreover, we provide fine-grained quantitative and visual attribute importance for each modality. Conclusion Our metric offers valuable insights that can support the advancement of multimodal model development and dataset creation. By introducing this method, we contribute to the growing field of interpretability in deep learning for multimodal research. This approach helps to facilitate the integration of multimodal AI into clinical practice. Our code is publicly available at https://github.com/ChristianGappGit/MC_MMD.

多模态可解释性医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。