通过人格情感对齐提升多模态情感分析精度
PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
- 在特征提取阶段引入人格特质,实现个性化情感嵌入
- 多层级融合策略逐步整合文本、视觉与音频情感信息
- 在两个主流数据集上达到当前最优性能,适合情感计算研究者
多模态情感分析(MSA)通过结合文本、视觉和音频模态识别人类情感。主要挑战在于如何有效融合不同模态的情感信息,这一问题通常出现在单模态特征提取和多模态特征融合阶段。现有方法在特征提取阶段仅获取浅层信息,忽视了不同人格下情感表达的差异;在融合阶段则直接拼接各模态特征,未考虑特征层面的异质性,最终影响模型识别效果。为此,我们提出一种人格-情感对齐的多层级融合框架(PSA-MF)。在特征提取阶段引入人格特质,首次提出人格-情感对齐方法,从文本模态生成个性化情感嵌入。在融合阶段,采用多层级融合策略,通过多模态预融合与增强融合逐步整合文本、视觉和音频模态的情感信息。该方法在两个常用数据集上进行了多组实验,取得了当前最优结果。
原文摘要 · Abstract (English)
Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction phase and the multimodal feature fusion phase. Existing methods extract only shallow information from unimodal features during the extraction phase, neglecting sentimental differences across different personalities. During the fusion phase, they directly merge the feature information from each modality without considering differences at the feature level. This ultimately affects the model's recognition performance. To address this problem, we propose a personality-sentiment aligned multi-level fusion framework. We introduce personality traits during the feature extraction phase and propose a novel personality-sentiment alignment method to obtain personalized sentiment embeddings from the textual modality for the first time. In the fusion phase, we introduce a novel multi-level fusion method. This method gradually integrates sentimental information from textual, visual, and audio modalities through multimodal pre-fusion and a multi-level enhanced fusion strategy. Our method has been evaluated through multiple experiments on two commonly used datasets, achieving state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。