arXiv:2608.16201cs.LG2026-08中稿 · NLPCC 2026

让大模型更懂多模态情感,通过分层次捕捉音频视频变化

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

论文配图:Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis
图 1 · 摘自论文原文
  • 分短、中、长时序捕捉音视频情感变化,保留细节与趋势
  • 用文本引导对齐音视频特征,提升模糊和中性样本识别率
  • 将多模态信息压缩成少量伪令牌,高效适配冻结大模型

多模态情感分析(MSA)旨在从文本、音频、视觉等异构输入中预测情感极性和强度。尽管大语言模型(LLMs)具备强大的语义先验,但如何有效融合音频和视觉信号仍具挑战。关键难点在于音频和视觉情感线索随不同时间尺度演变,而现有基于LLM的方法常通过浅层投影或粗粒度池化压缩这些信号,导致跨模态对齐减弱并丢失细粒度情感信息。本文提出MGSI:一种面向LLM的多粒度情感融合框架。MGSI首先在短、中、长三个时间尺度上编码音频和视觉流,保留局部波动与全局情感趋势;随后通过文本引导对齐优化非文本特征,并引入极性与强度感知增强机制,更好处理模糊及近中性样本;最终将融合后的多模态表示压缩为少量伪令牌,以高效条件化冻结的LLM。在四个公开基准上的实验表明,MGSI显著优于冻结LLM基线,且与强大多模态方法相当。消融与敏感性分析进一步验证了多粒度时序建模、文本引导精炼及自适应情感校准的有效性。

原文摘要 · Abstract (English)

Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. While large language models (LLMs) offer strong semantic priors for MSA, effectively incorporating audio and visual signals effectively remains challenging. A key challenge is that audio and visual sentiment cues evolve over different temporal scales, yet many LLM-based methods compress these signals through shallow projection or coarse pooling before fusing them with text, which can weaken cross-modal alignment and erase fine-grained affective information. We propose MGSI, a multi-granularity sentiment integration framework for LLM-based MSA. MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends. It then refines non-text features through text-guided alignment, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples. The resulting multimodal representation is finally compressed into a small set of pseudo-tokens for efficient conditioning of a frozen LLM. Experiments on four public benchmarks show that MGSI substantially outperforms frozen-LLM baselines and remains competitive with strong multimodal methods. Further ablation and sensitivity analyses support the effectiveness of multi-granularity temporal modeling, text-guided refinement, and adaptive sentiment calibration.

多模态情感分析大模型时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。