通过初始化调控缓解多模态融合中的模态竞争问题
Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition
- 先单模态训练再联合训练,提前调整初始状态
- 提出快速部分信息分解法,精准量化模态间关系
- 在多个数据集上达到领先性能,适合多模态研究者
多模态融合在联合训练中常遭遇模态竞争,某一模态主导学习过程,导致其他模态优化不足。现有方法多聚焦于联合训练阶段,忽视模型初始状态的关键影响。本文提出两阶段训练框架:先通过单模态训练塑造初始状态。首先定义有效竞争强度(ECS)量化模态竞争能力,理论分析表明合理设定初始ECS可获得更紧的误差界。但ECS在深度网络中计算不可行。为此,我们构建包含两个核心组件的框架:细粒度可计算诊断指标与异步训练控制器。针对指标,证明互信息(MI)是ECS的合理代理;进一步提出快速部分信息分解(FastPID),一种高效可微的求解器,将联合分布信息分解为模态特有、冗余与协同三部分。基于此,异步控制器监控模态特有性,定位协同峰值以确定最优联合训练起始点。在多种基准测试中,本方法表现达当前最优。研究证实,预先塑造融合前模型的初始状态是一种有效策略,可在竞争发生前即实现协同融合。
原文摘要 · Abstract (English)
Multi-modal fusion often suffers from modality competition during joint training, where one modality dominates the learning process, leaving others under-optimized. Overlooking the critical impact of the model's initial state, most existing methods address this issue during the joint learning stage. In this study, we introduce a two-stage training framework to shape the initial states through unimodal training before the joint training. First, we propose the concept of Effective Competitive Strength (ECS) to quantify a modality's competitive strength. Our theoretical analysis further reveals that properly shaping the initial ECS by unimodal training achieves a provably tighter error bound. However, ECS is computationally intractable in deep neural networks. To bridge this gap, we develop a framework comprising two core components: a fine-grained computable diagnostic metric and an asynchronous training controller. For the metric, we first prove that mutual information(MI) is a principled proxy for ECS. Considering MI is induced by per-modality marginals and thus treats each modality in isolation, we further propose FastPID, a computationally efficient and differentiable solver for partial information decomposition, which decomposes the joint distribution's information into fine-grained measurements: modality-specific uniqueness, redundancy, and synergy. Guided by these measurements, our asynchronous controller dynamically balances modalities by monitoring uniqueness and locates the ideal initial state to start joint training by tracking peak synergy. Experiments on diverse benchmarks demonstrate that our method achieves state-of-the-art performance. Our work establishes that shaping the pre-fusion models' initial state is a powerful strategy that eases competition before it starts, reliably unlocking synergistic multi-modal fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。