提出对称互促融合机制,提升多模态情感分析的特征互补性。
Multimodal Sentiment Analysis based on Multi-channel and Symmetric Mutual Promotion Feature Fusion
- 多通道提取视觉与听觉特征,增强模态内表征能力。
- 对称互促融合使跨模态信息交换更充分,准确率显著提升。
- 适合需要精准情感识别的应用场景,如人机交互系统。
多模态情感分析是人机交互与情感计算领域关键技术。准确识别人类情绪状态对促进人机顺畅沟通至关重要。尽管已有一定进展,仍面临两大挑战:单一模态数据提取特征有限且不充分;多数研究仅关注模态间特征一致性,忽视差异性,导致特征融合不足。本文首先通过多通道方法提取更全面的特征信息,在视觉与听觉模态中采用双通道设计以增强模态内表示。其次,提出对称互促(SMP)跨模态融合方法,结合对称交叉注意力与自注意力机制,前者捕捉其他模态有用信息,后者建模上下文依赖关系,促进模态间有效信息交换。最后,融合模态内特征与跨模态融合特征,兼顾模态间互补性与差异性。在两个基准数据集上的实验验证了该方法的有效性与优越性。
原文摘要 · Abstract (English)
Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and machines. Despite some progress in multimodal sentiment analysis research, numerous challenges remain. The first challenge is the limited and insufficiently rich features extracted from single modality data. Secondly, most studies focus only on the consistency of inter-modal feature information, neglecting the differences between features, resulting in inadequate feature information fusion. In this paper, we first extract multi-channel features to obtain more comprehensive feature information. We employ dual-channel features in both the visual and auditory modalities to enhance intra-modal feature representation. Secondly, we propose a symmetric mutual promotion (SMP) inter-modal feature fusion method. This method combines symmetric cross-modal attention mechanisms and self-attention mechanisms, where the cross-modal attention mechanism captures useful information from other modalities, and the self-attention mechanism models contextual information. This approach promotes the exchange of useful information between modalities, thereby strengthening inter-modal interactions. Furthermore, we integrate intra-modal features and inter-modal fused features, fully leveraging the complementarity of inter-modal feature information while considering feature information differences. Experiments conducted on two benchmark datasets demonstrate the effectiveness and superiority of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。