动态选择主模态并压缩冗余信息,提升视频情感分析效果
Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
- 用图结构压缩声画模态序列冗余,减少噪声干扰
- 根据样本自适应选择主模态,避免固定策略失效
- 在主流数据集上优于现有方法,尤其适合模态不平衡场景
多模态情感分析(MSA)旨在从视频的语言、声音和视觉数据中预测情感。然而,单模态性能不均衡常导致融合表示不佳。现有方法通常采用固定主模态策略以强化主导模态优势,但难以适应不同样本中模态重要性的动态变化。此外,非语言模态存在序列冗余和噪声问题,当其作为主输入时会降低模型表现。为此,本文提出模态优化与动态主模态选择框架(MODS)。首先构建基于图的动态序列压缩器(GDC),利用胶囊网络与图卷积减少声/视频模态的序列冗余;其次设计样本自适应主模态选择器(MSelector),实现动态主导性判断;最后提出主模态中心交叉注意力模块(PCCA),增强主导模态同时促进跨模态交互。在四个基准数据集上的大量实验表明,MODS显著优于当前最优方法,有效平衡模态贡献并消除冗余噪声。
原文摘要 · Abstract (English)
Multimodal Sentiment Analysis (MSA) aims to predict sentiment from language, acoustic, and visual data in videos. However, imbalanced unimodal performance often leads to suboptimal fused representations. Existing approaches typically adopt fixed primary modality strategies to maximize dominant modality advantages, yet fail to adapt to dynamic variations in modality importance across different samples. Moreover, non-language modalities suffer from sequential redundancy and noise, degrading model performance when they serve as primary inputs. To address these issues, this paper proposes a modality optimization and dynamic primary modality selection framework (MODS). First, a Graph-based Dynamic Sequence Compressor (GDC) is constructed, which employs capsule networks and graph convolution to reduce sequential redundancy in acoustic/visual modalities. Then, we develop a sample-adaptive Primary Modality Selector (MSelector) for dynamic dominance determination. Finally, a Primary-modality-Centric Cross-Attention (PCCA) module is designed to enhance dominant modalities while facilitating cross-modal interaction. Extensive experiments on four benchmark datasets demonstrate that MODS outperforms state-of-the-art methods, achieving superior performance by effectively balancing modality contributions and eliminating redundant noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。