分离三类语义空间,提升多模态情感分析精度
Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
- 将特征分解为全局、成对共享和私有三类子空间
- 在CMU-MOSI上达0.691的MAE,CMU-MOSEI上54.9%准确率
- 适合需要精细跨模态融合的多模态任务研究者
多模态情感分析(MSA)融合语言、视觉和听觉模态以推断人类情感。现有方法多关注全局共享表示或单模态特异性特征,忽视仅由特定模态对共享的信号,限制了表征的表达力与判别性。为此,我们提出三子空间解耦(TSD)框架,将特征显式分解为三类互补子空间:全局一致性子空间、子模态共享子空间(建模成对跨模态协同)和私有子空间(保留模态特异性线索)。为保证子空间纯净独立,引入解耦监督器及结构化正则化损失。进一步设计子空间感知交叉注意力(SACA)融合模块,自适应整合三类子空间信息,获得更丰富鲁棒的表征。在CMU-MOSI和CMU-MOSEI上的实验表明,TSD在所有关键指标上均达当前最优性能,分别取得0.691 MAE和54.9% ACC-7,且在多模态意图识别任务中具有良好迁移能力。消融实验证实,三子空间解耦与SACA共同增强对多粒度跨模态情感线索的建模。
原文摘要 · Abstract (English)
Multimodal Sentiment Analysis (MSA) integrates language, visual, and acoustic modalities to infer human sentiment. Most existing methods either focus on globally shared representations or modality-specific features, while overlooking signals that are shared only by certain modality pairs. This limits the expressiveness and discriminative power of multimodal representations. To address this limitation, we propose a Tri-Subspace Disentanglement (TSD) framework that explicitly factorizes features into three complementary subspaces: a common subspace capturing global consistency, submodally-shared subspaces modeling pairwise cross-modal synergies, and private subspaces preserving modality-specific cues. To keep these subspaces pure and independent, we introduce a decoupling supervisor together with structured regularization losses. We further design a Subspace-Aware Cross-Attention (SACA) fusion module that adaptively models and integrates information from the three subspaces to obtain richer and more robust representations. Experiments on CMU-MOSI and CMU-MOSEI demonstrate that TSD achieves state-of-the-art performance across all key metrics, reaching 0.691 MAE on CMU-MOSI and 54.9% ACC-7 on CMU-MOSEI, and also transfers well to multimodal intent recognition tasks. Ablation studies confirm that tri-subspace disentanglement and SACA jointly enhance the modeling of multi-granular cross-modal sentiment cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。