用信息论判断跨模态知识蒸馏是否有效,选对老师能提升模型表现。
Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning
- 基于互信息设计理论准则,判断跨模态蒸馏是否有效。
- 实证验证在图像、文本、视频等多类数据上均适用。
- 适合做多模态融合或想优化弱模态性能的研究者参考。
多模态数据的快速增长推动了跨模态知识蒸馏(KD)技术的发展,即通过更丰富的“教师”模态向较弱的“学生”模态传递信息以提升性能。然而,尽管广泛应用,跨模态KD并不总能带来提升,主要原因在于缺乏理论指导。为此,我们提出跨模态互补性假说(CCH):当教师与学生表示之间的互信息超过学生表示与标签之间的互信息时,跨模态KD才有效。我们在联合高斯模型中理论验证了该假说,并在包括图像、文本、视频、音频及癌症组学数据在内的多种多模态数据集上进行了实证检验。本研究建立了理解跨模态KD的新理论框架,并提供了基于CCH准则选择最优教师模态的实践指南。
原文摘要 · Abstract (English)
The rapid increase in multimodal data availability has sparked significant interest in cross-modal knowledge distillation (KD) techniques, where richer "teacher" modalities transfer information to weaker "student" modalities during model training to improve performance. However, despite successes across various applications, cross-modal KD does not always result in improved outcomes, primarily due to a limited theoretical understanding that could inform practice. To address this gap, we introduce the Cross-modal Complementarity Hypothesis (CCH): we propose that cross-modal KD is effective when the mutual information between teacher and student representations exceeds the mutual information between the student representation and the labels. We theoretically validate the CCH in a joint Gaussian model and further confirm it empirically across diverse multimodal datasets, including image, text, video, audio, and cancer-related omics data. Our study establishes a novel theoretical framework for understanding cross-modal KD and offers practical guidelines based on the CCH criterion to select optimal teacher modalities for improving the performance of weaker modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。