通过动态调整模态混合比例,提升多模态视频理解的泛化能力
Mixup Helps Understanding Multimodal Video Better
- 在融合特征层应用混合法,生成虚拟样本增强模型鲁棒性
- 相比传统方法,显著提升弱模态贡献,改善跨数据集性能
- 适合关注多模态对齐与模型公平性的研究者参考
多模态视频理解在动作识别和情绪分类等任务中至关重要,通过融合不同模态信息实现。然而,多模态模型易过度依赖强模态,导致弱模态贡献被压制。为此,我们提出多模态混合法(MM),在聚合的多模态特征层面应用混合法,生成虚拟特征-标签对以缓解过拟合。尽管MM有效提升了泛化能力,但其对各模态处理均等,未考虑训练中的模态不平衡问题。在此基础上,我们进一步提出平衡式多模态混合法(B-MM),根据各模态对学习目标的相对贡献动态调整混合比例。在多个数据集上的大量实验表明,该方法能有效提升泛化性和多模态鲁棒性。
原文摘要 · Abstract (English)
Multimodal video understanding plays a crucial role in tasks such as action recognition and emotion classification by combining information from different modalities. However, multimodal models are prone to overfitting strong modalities, which can dominate learning and suppress the contributions of weaker ones. To address this challenge, we first propose Multimodal Mixup (MM), which applies the Mixup strategy at the aggregated multimodal feature level to mitigate overfitting by generating virtual feature-label pairs. While MM effectively improves generalization, it treats all modalities uniformly and does not account for modality imbalance during training. Building on MM, we further introduce Balanced Multimodal Mixup (B-MM), which dynamically adjusts the mixing ratios for each modality based on their relative contributions to the learning objective. Extensive experiments on several datasets demonstrate the effectiveness of our methods in improving generalization and multimodal robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。