arXiv:2412.10460cs.CVcs.AI2024-12AAAI被引 84

用文字描述视频音频情绪,提升多模态情感分析精度。

Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content

  • 将音视频内容转为情绪文字描述,增强情感特征。
  • 在MOSI、MOSEI等数据集上超越现有模型表现。
  • 适合关注细微情感变化的多模态研究者。

多模态情感分析(MSA)旨在融合文本、音频和视觉数据以全面解析人类情绪。然而,在音频与视频表达中识别微妙情绪差异仍具挑战性,尤其当不同片段情感极性相似时。本文提出DEVA,一种基于文本情感描述的渐进式融合框架,聚焦音视频模态中的情绪相关属性,促进多模态融合。DEVA通过情感描述生成器(EDG)将原始音视频数据转化为文本化情绪描述,强化其情感特征,并与源数据融合生成更丰富特征。同时引入文本引导渐进融合模块(TPF),以不同层级的文本作为核心引导,逐步融合视觉与音频子模态,缓解文本与音视频间的差异。在MOSI、MOSEI和CH-SIMS等多个基准数据集上的实验表明,该方法显著优于当前最优模型。细粒度情绪实验进一步验证了DEVA对细微情感变化的强敏感性。

原文摘要 · Abstract (English)

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video expressions poses a formidable challenge, particularly when emotional polarities across various segments appear similar. In this paper, our objective is to spotlight emotion-relevant attributes of audio and visual modalities to facilitate multimodal fusion in the context of nuanced emotional shifts in visual-audio scenarios. To this end, we introduce DEVA, a progressive fusion framework founded on textual sentiment descriptions aimed at accentuating emotional features of visual-audio content. DEVA employs an Emotional Description Generator (EDG) to transmute raw audio and visual data into textualized sentiment descriptions, thereby amplifying their emotional characteristics. These descriptions are then integrated with the source data to yield richer, enhanced features. Furthermore, DEVA incorporates the Text-guided Progressive Fusion Module (TPF), leveraging varying levels of text as a core modality guide. This module progressively fuses visual-audio minor modalities to alleviate disparities between text and visual-audio modalities. Experimental results on widely used sentiment analysis benchmark datasets, including MOSI, MOSEI, and CH-SIMS, underscore significant enhancements compared to state-of-the-art models. Moreover, fine-grained emotion experiments corroborate the robust sensitivity of DEVA to subtle emotional variations.

多模态情感分析文本生成融合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。