通过对齐多模态表示提升情感分析效果,让视觉信息变成可理解的文本。
Explicit Representation Alignment for Multimodal Sentiment Analysis

- 用视觉语言模型将图像转为结构化文本,统一到共同语义空间。
- 在多个数据集上超越强基线,实现当前最佳性能。
- 适合关注多模态对齐与可解释性推理的研究者。
多模态情感分析旨在通过联合建模文本、图像等异构模态来理解人类情感与情绪。然而,现有模型常无法持续优于仅依赖文本的基线,且融合策略表现差异显著。本文识别出独立预训练模态编码器间表征错位是有效多模态学习的关键瓶颈,并通过受控实验表明:融合前的对齐比融合复杂度更为重要。为此,我们提出统一的多模态情感分析框架,利用视觉语言模型(VLMs)将视觉内容转换为结构化文本描述,将异构模态投影至共享语言空间,支持可解释的以文本为中心的推理。为进一步提升鲁棒性,引入混合学习策略,结合语义标记选择与批次级均匀性正则化目标,促使全局特征空间更分散稳定,同时缓解VLM生成描述带来的噪声。在多个多模态情感与情绪基准测试上,该方法持续优于强单模态与多模态基线,达到当前最优性能。分析进一步凸显了表征对齐在多模态情感学习中的关键作用。
原文摘要 · Abstract (English)
Multimodal affective analysis aims to understand human sentiment and emotion by jointly modeling heterogeneous modalities such as text and images. However, multimodal models often fail to consistently outperform strong text-only baselines, with performance varying significantly across fusion strategies. In this work, we identify representation misalignment between independently pretrained modality encoders as a key bottleneck for effective multimodal learning, and show through controlled experiments that alignment prior to fusion is often more important than fusion complexity. To address this issue, we propose a unified multimodal affective analysis framework that leverages vision-language models (VLMs) to convert visual content into structured textual descriptions, projecting heterogeneous modalities into a shared linguistic space and enabling interpretable text-centric reasoning. To further improve robustness, we introduce a hybrid learning strategy that combines semantic token selection with a batch-level uniformity regularization objective, encouraging a more dispersed and stable global feature space while mitigating noise introduced by VLM-generated descriptions. Experiments on multiple multimodal sentiment and emotion benchmarks show that our method consistently outperforms strong unimodal and multimodal baselines, achieving state-of-the-art performance. Our analysis further highlights the critical role of representation alignment in multimodal affective learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。