arXiv:2608.24947cs.LGcs.AI2026-08中稿 · Neurocomputing Jou…

解决多模态模型训练中的梯度不稳问题,提升融合效果与稳定性。

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

论文配图:CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery
图 1 · 摘自论文原文
  • 通过校准门控机制与融合手术,动态调控多模态梯度传播。
  • 在多个数据集上超越或持平强基线模型,减少冲突梯度。
  • 无需修改结构,适配各类多模态任务,尤其适合复杂融合场景。

端到端训练多模态神经网络常因三种耦合故障模式导致学习不稳定:(i) 模态失衡,某分支主导梯度优化;(ii) 门控不稳,噪声置信信号引发异常模态选择;(iii) 融合干扰,共享融合层处模态特异性梯度冲突。本文提出CAT-GS(校准、自适应、阈值门控融合手术),一种基于神经动力学的优化控制器,无需修改模型架构、融合模块或任务损失,在反向传播中运行。通过温度缩放与指数移动平均平滑的教师可靠性校准,采用边际阈值策略,在预热丢弃、弱模态优先和弱偏融合间切换;通过截断梯度预算归一化稳定激进门控下的梯度幅值;并应用仅融合版PCGrad,减少主共享瓶颈处破坏性跨模态干扰。在音频-视觉模式识别基准(CREMA-D, AV-MNIST, VGGSound)、三模态设置(UR-FUNNY)、可控合成数据(CG-MNIST)及跨域基准(AVE, CMU-MOSI)上评估,CAT-GS在多种设置下优于或匹配强失衡感知基线(包括OGM-GE、G²D、UMT),且呈现更平滑的门控行为与更少冲突的融合梯度。

原文摘要 · Abstract (English)

End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-specific gradients conflict at the shared fusion layer. We propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing applications. CAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses. Through calibration of teacher-derived reliability via temperature scaling and EMA smoothing, CAT-GS stabilizes neural dynamics using a margin-thresholded policy to switch between warm-up dropout, weak-modality prioritization, and weak-biased blending, stabilizes gradient magnitudes under aggressive gating via capped gradient-budget renormalization, and applies fusion-only PCGrad to reduce destructive cross-modal interference at the primary shared bottleneck. We evaluate CAT-GS on audio--visual multimodal pattern recognition benchmarks (CREMA-D, AV-MNIST, and VGGSound), a tri-modal setting (UR-FUNNY), controlled synthetic data (CG-MNIST), and additional cross-domain benchmarks (AVE and CMU-MOSI). CAT-GS improves or matches fused multimodal accuracy against strong imbalance-aware baselines (including OGM-GE, G$^2$D, and UMT) across settings, and yields smoother gating behavior with fewer conflicting fusion gradients.

多模态梯度优化融合机制稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。