提出轻量级方法,精准量化多模态数据中每条样本的交互关系。
Efficient Quantification of Multimodal Interaction at Sample Level
- 基于点熵理论设计高效估计器,实现连续分布下样本级交互量化。
- 在合成与真实数据上验证精度,计算效率优于现有方法。
- 可支持知识蒸馏、模型融合等应用,适合多模态系统分析场景。
多模态间的冗余、独特性和协同作用共同决定信息构成。理解这些交互对分析多模态系统中的信息动态至关重要,但其在样本层面的精确量化面临重大理论与计算挑战。为此,我们提出轻量级样本级多模态交互(LSMI)估计器,严格基于点熵信息论。首先构建冗余估计框架,采用合适的点熵度量来量化最可分解和可测量的交互。在此基础上,提出一种通用交互估计方法,结合高效熵估计,专门针对连续分布下的样本级估计。在合成与真实数据集上的大量实验验证了LSMI的高精度与高效率。关键的是,该样本级方法揭示了多模态数据中细粒度的样本级与类别级动态,支持冗余引导的样本划分、针对性知识蒸馏及交互感知的模型集成等实际应用。代码已开源:https://github.com/GeWu-Lab/LSMI_Estimator。
原文摘要 · Abstract (English)
Interactions between modalities -- redundancy, uniqueness, and synergy -- collectively determine the composition of multimodal information. Understanding these interactions is crucial for analyzing information dynamics in multimodal systems, yet their accurate sample-level quantification presents significant theoretical and computational challenges. To address this, we introduce the Lightweight Sample-wise Multimodal Interaction (LSMI) estimator, rigorously grounded in pointwise information theory. We first develop a redundancy estimation framework, employing an appropriate pointwise information measure to quantify this most decomposable and measurable interaction. Building upon this, we propose a general interaction estimation method that employs efficient entropy estimation, specifically tailored for sample-wise estimation in continuous distributions. Extensive experiments on synthetic and real-world datasets validate LSMI's precision and efficiency. Crucially, our sample-wise approach reveals fine-grained sample- and category-level dynamics within multimodal data, enabling practical applications such as redundancy-informed sample partitioning, targeted knowledge distillation, and interaction-aware model ensembling. The code is available at https://github.com/GeWu-Lab/LSMI_Estimator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。