arXiv:2508.19769cs.CV2025-08被引 2

提出新方法让强弱模态共同提升,不压制任何一方。

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

  • 根据网络内部优化状态动态调节各层模态学习强度
  • 在多个数据集上超越现有最强方法,提升明显
  • 适合需要平衡多模态训练的场景,如跨模态理解

多模态学习虽显著提升性能,但仍面临模态不平衡问题。现有方法通常通过抑制主导模态来促进弱模态,但会损害整体表现。本文分析发现其根源在于网络内优化偏差。为此提出自适应网络内调制(AIM),首次实现不压制任一模态的均衡学习。AIM将主导模态中未充分优化的参数解耦至辅助模块,引导其与弱模态联合训练;同时按网络深度评估模态不平衡程度,自适应调整各层调制强度。实验表明,AIM在多个基准上优于当前最优方法,且对不同主干网络、融合策略和优化器均具强泛化性。

原文摘要 · Abstract (English)

Multimodal learning has significantly enhanced machine learning performance but still faces numerous challenges and limitations. Imbalanced multimodal learning is one of the problems extensively studied in recent works and is typically mitigated by modulating the learning of each modality. However, we find that these methods typically hinder the dominant modality's learning to promote weaker modalities, which affects overall multimodal performance. We analyze the cause of this issue and highlight a commonly overlooked problem: optimization bias within networks. To address this, we propose Adaptive Intra-Network Modulation (AIM) to improve balanced modality learning. AIM accounts for differences in optimization state across parameters and depths within the network during modulation, achieving balanced multimodal learning without hindering either dominant or weak modalities for the first time. Specifically, AIM decouples the dominant modality's under-optimized parameters into Auxiliary Blocks and encourages reliance on these performance-degraded blocks for joint training with weaker modalities. This approach effectively prevents suppression of weaker modalities while enabling targeted optimization of under-optimized parameters to improve the dominant modality. Additionally, AIM assesses modality imbalance level across network depths and adaptively adjusts modulation strength at each depth. Experimental results demonstrate that AIM outperforms state-of-the-art imbalanced modality learning methods across multiple benchmarks and exhibits strong generalizability across different backbones, fusion strategies, and optimizers.

多模态学习模态平衡自适应调制神经网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。