通过双向引导的稳定机制,提升多模态情感分析的时序与跨模态一致性。
Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization
- 引入跨模态不变融合,增强不同模态间的稳定表示
- 设计序列变化正则化,约束反向传播中的学习轨迹
- 在三个公开数据集上验证效果,显著提升情感预测稳定性
在多模态情感分析中,实现不同模态间的一致情感表征仍面临挑战。然而,时间上的快速情绪波动常导致表征不稳定,影响预测性能。为此,我们提出一种鲁棒的情感表征双重增强策略,同时优化时序与模态维度,通过前向与反向传播中的定向机制实现。在模态维度,引入模态不变融合机制,旨在捕捉跨模态共享的共性与稳定表征;在时序维度,施加专门的序列变化正则化项,在反向传播中调节模型学习路径,本质为一维线性差分形式的总变差正则化。在三个标准公开数据集上的大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Achieving consistent sentiment representation across diverse modalities remains a key challenge in multimodal sentiment analysis. However, rapid emotional fluctuations over time often introduce instability, leading to compromised prediction performance. To address this challenge, we propose a robust sentiment representation dual enhancement strategy that simultaneously enhances the temporal and modality dimensions, guided by targeted mechanisms in both forward and backward propagation. Specifically, in the modality dimension, we introduce a modality invariant fusion mechanism that fosters stable cross-modal representations, which aim to capture the common and stable representations shared across different modalities. In the temporal dimension, we impose a specialized sequential variation regularization term that regulates the model's learning trajectory during backward propagation, which is essentially total variation regularization degenerated into one-dimensional linear differences. Extensive experiments on three standard public datasets validate the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。