先解耦时序与空间信息,提升多模态情感分析效果
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
- 先分离每模态的时序与空间特征,再跨模态对齐
- 在CMU-MOSEI数据集上准确率提升2.1个百分点
- 适合需要可解释性的多模态情感分析场景
多模态情感分析融合语言、视觉和听觉信息。主流方法基于模态不变与特定因子分解或复杂融合,仍依赖时空混合建模,忽视时空异质性,导致信息不对称,性能受限。为此,提出TSDA(时序-空间解耦先行),在模态交互前显式将各模态解耦为时序动态与空间结构上下文。每模态通过时序编码器与空间编码器分别投影至独立的时序与空间表征。因子一致跨模态对齐仅在同类型特征间进行:时序对时序,空间对空间。因子特异性监督与去相关正则化减少跨因子泄露,同时保留互补性。门控重耦模块随后对齐后的流进行重组以完成任务。大量实验表明,TSDA优于基线模型。消融分析证实设计的必要性与可解释性。
原文摘要 · Abstract (English)
Multimodal Sentiment Analysis integrates Linguistic, Visual, and Acoustic. Mainstream approaches based on modality-invariant and modality-specific factorization or on complex fusion still rely on spatiotemporal mixed modeling. This ignores spatiotemporal heterogeneity, leading to spatiotemporal information asymmetry and thus limited performance. Hence, we propose TSDA, Temporal-Spatial Decouple before Act, which explicitly decouples each modality into temporal dynamics and spatial structural context before any interaction. For every modality, a temporal encoder and a spatial encoder project signals into separate temporal and spatial body. Factor-Consistent Cross-Modal Alignment then aligns temporal features only with their temporal counterparts across modalities, and spatial features only with their spatial counterparts. Factor specific supervision and decorrelation regularization reduce cross factor leakage while preserving complementarity. A Gated Recouple module subsequently recouples the aligned streams for task. Extensive experiments show that TSDA outperforms baselines. Ablation analysis studies confirm the necessity and interpretability of the design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。