将多模态情感表示拆分为共性与特异性成分,提升情感理解准确率。
Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing
- 通过分解视觉与文本特征,分离出跨模态共性与模态特有信息
- 在三大情感任务上均超越现有模型,最高提升4.2%准确率
- 适合需要精细情感分析的智能交互系统开发者
多模态情感计算旨在从图像、文本等多元数据中自动识别和理解人类态度,以增强人机交互与情绪认知。现有方法多依赖单模态分析或简单的跨模态融合,难以捕捉不同模态间复杂且矛盾的证据。本文提出一种基于大语言模型(LLM)的新方法,显式将视觉与文本表征分解为共享(模态不变)和模态特有两部分。具体流程:首先使用预训练多模态编码器对输入模态进行编码与对齐,再通过表示分解框架分离出共有的情感内容与独特的线索,最后利用注意力机制整合这些分解后的信号,生成动态软提示输入多模态LLM。在三个代表性任务上——多模态细粒度情感分析、多模态情绪分析、仇恨表情包检测——的大量实验表明,该方法持续优于强基线及当前最优模型。
原文摘要 · Abstract (English)
Multi-modal affective computing aims to automatically recognize and interpret human attitudes from diverse data sources such as images and text, thereby enhancing human-computer interaction and emotion understanding. Existing approaches typically rely on unimodal analysis or straightforward fusion of cross-modal information that fail to capture complex and conflicting evidence presented across different modalities. In this paper, we propose a novel LLM-based approach for affective computing that explicitly deconstructs visual and textual representations into shared (modality-invariant) and modality-specific components. Specifically, our approach firstly encodes and aligns input modalities using pre-trained multi-modal encoders, then employs a representation decomposition framework to separate common emotional content from unique cues, and finally integrates these decomposed signals via an attention mechanism to form a dynamic soft prompt for a multi-modal LLM. Extensive experiments on three representative tasks for affective computing, namely, multi-modal aspect-based sentiment analysis, multi-modal emotion analysis, and hateful meme detection, demonstrate the effectiveness of our approach, which consistently outperforms strong baselines and state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。