arXiv:2504.15848cs.CL2025-04中稿 · TAFFC 2025被引 47

通过认知与审美因果机制,提升多模态情感分析的细粒度理解能力。

Exploring Cognitive and Aesthetic Causality for Multimodal Aspect-Based Sentiment Analysis

  • 融合视觉块对齐与文本描述生成,捕捉图像中细粒度语义特征。
  • 利用大模型生成情感成因与印象,增强对情绪与认知共振的理解。
  • 在标准数据集上表现优于GPT-4o,适合需要深度情感推理的应用。

多模态方面级情感分类(MASC)因社交平台用户生成内容激增而成为新兴任务,旨在预测针对特定方面目标(如文本-图像对中明确提及的实体或属性)的情感极性。尽管现有研究取得显著进展,但在理解细粒度视觉内容及由语义内容和印象产生的认知逻辑(情感唤起的认知解释)方面仍存在明显差距。本研究提出Chimera框架,通过认知与审美因果理解机制,提取方面的细粒度整体特征,并从语义视角与情感-认知共鸣(情绪反应与认知解释的协同效应)中推断情感表达的根本动因。该框架首先引入视觉块特征进行块-词对齐;同时提取粗粒度视觉特征(如整体图像表征)与细粒度视觉区域(如与方面相关的区域),并将其转化为对应文本描述(如面部、美学特征)。最后,利用大语言模型(LLM)生成的情感成因与印象,增强模型对语义内容引发的情感线索及情感-认知共鸣的感知能力。在标准MASC数据集上的实验结果表明,所提模型有效提升了性能,且相较于GPT-4o等大模型展现出更强的灵活性。完整代码与数据集已公开于 https://github.com/Xillv/Chimera。

原文摘要 · Abstract (English)

Multimodal aspect-based sentiment classification (MASC) is an emerging task due to an increase in user-generated multimodal content on social platforms, aimed at predicting sentiment polarity toward specific aspect targets (i.e., entities or attributes explicitly mentioned in text-image pairs). Despite extensive efforts and significant achievements in existing MASC, substantial gaps remain in understanding fine-grained visual content and the cognitive rationales derived from semantic content and impressions (cognitive interpretations of emotions evoked by image content). In this study, we present Chimera: a cognitive and aesthetic sentiment causality understanding framework to derive fine-grained holistic features of aspects and infer the fundamental drivers of sentiment expression from both semantic perspectives and affective-cognitive resonance (the synergistic effect between emotional responses and cognitive interpretations). Specifically, this framework first incorporates visual patch features for patch-word alignment. Meanwhile, it extracts coarse-grained visual features (e.g., overall image representation) and fine-grained visual regions (e.g., aspect-related regions) and translates them into corresponding textual descriptions (e.g., facial, aesthetic). Finally, we leverage the sentimental causes and impressions generated by a large language model (LLM) to enhance the model's awareness of sentimental cues evoked by semantic content and affective-cognitive resonance. Experimental results on standard MASC datasets demonstrate the effectiveness of the proposed model, which also exhibits greater flexibility to MASC compared to LLMs such as GPT-4o. We have publicly released the complete implementation and dataset at https://github.com/Xillv/Chimera

多模态情感分析认知机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。