arXiv:2508.18551cs.LG2025-08EMNLP被引 1

解决多模态模型中噪声模态干扰问题,动态调整各模态权重。

BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration

  • 基于实例级KL散度和模态级互信息,非参数化动态加权
  • 在情感回归与临床分类任务中显著提升性能
  • 无需额外参数,可扩展至任意数量模态,适合高噪声场景

混合专家(MoE)模型通过在不同模态间实现模块化专精,增强了多模态学习能力。然而,当新增模态引入的噪声超过互补信息时,其有效性仍不明确。现有方法如部分信息分解难以扩展至两模态以上,且缺乏实例级控制粒度。我们提出超越两模态加权(BTW),一种双层非参数加权框架,结合实例级Kullback-Leibler(KL)散度与模态级互信息(MI),在训练过程中动态调节模态重要性。该方法不需额外参数,可应用于任意数量模态。具体而言,BTW通过测量每个单模态与当前多模态预测间的分布差异计算实例级KL权重,并通过估计单模态与多模态输出间的全局对齐程度获得模态级MI权重。在情感回归与临床分类任务上的大量实验表明,该方法显著提升了回归性能与多类分类准确率。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) models have become increasingly powerful in multimodal learning by enabling modular specialization across modalities. However, their effectiveness remains unclear when additional modalities introduce more noise than complementary information. Existing approaches, such as the Partial Information Decomposition, struggle to scale beyond two modalities and lack the resolution needed for instance-level control. We propose Beyond Two-modality Weighting (BTW), a bi-level, non-parametric weighting framework that combines instance-level Kullback-Leibler (KL) divergence and modality-level mutual information (MI) to dynamically adjust modality importance during training. Our method does not require additional parameters and can be applied to an arbitrary number of modalities. Specifically, BTW computes per-example KL weights by measuring the divergence between each unimodal and the current multimodal prediction, and modality-wide MI weights by estimating global alignment between unimodal and multimodal outputs. Extensive experiments on sentiment regression and clinical classification demonstrate that our method significantly improves regression performance and multiclass classification accuracy.

多模态加权机制非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。