按顺序学习文本和视频特征,提升跨域情感分析效果
Learning in Order! A Sequential Strategy to Learn Invariant Features for Multimodal Sentiment Analysis
- 先学文本的域不变特征,再用其指导视频的稀疏通用特征学习
- 在单源和多源设置下均显著优于当前最佳方法
- 选择独立且与情感标签强相关的特征,适合跨域情感分析研究
本文提出一种新颖且简单的顺序学习策略,用于在视频和文本上训练多模态情感分析模型。为在未见的分布外数据上估计情感极性,该模型在单一源域或多源域下使用此策略进行训练。策略首先从文本中学习域不变特征,随后在文本特征的辅助下,学习视频的稀疏域无关特征。实验结果表明,该模型在单源和多源设置下平均性能显著优于现有最先进方法。特征选择过程优先选取彼此独立且与情感标签强相关的特征。为促进该领域研究,代码将在论文接受后公开。
原文摘要 · Abstract (English)
This work proposes a novel and simple sequential learning strategy to train models on videos and texts for multimodal sentiment analysis. To estimate sentiment polarities on unseen out-of-distribution data, we introduce a multimodal model that is trained either in a single source domain or multiple source domains using our learning strategy. This strategy starts with learning domain invariant features from text, followed by learning sparse domain-agnostic features from videos, assisted by the selected features learned in text. Our experimental results demonstrate that our model achieves significantly better performance than the state-of-the-art approaches on average in both single-source and multi-source settings. Our feature selection procedure favors the features that are independent to each other and are strongly correlated with their polarity labels. To facilitate research on this topic, the source code of this work will be publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。