arXiv:2409.00304cs.CV2024-09IJCV被引 18

让大模型更懂视频如何打动人心,提升情感推理能力。

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

  • 通过帧级与令牌级双重注意力机制,聚焦情绪触发区域。
  • 在多个数据集上显著提升情感反应预测准确率,解释更合理。
  • 适合需要理解观众情绪的视频内容评估与社交智能系统开发。

预测和推理视频如何影响人类情感,对构建具有社会智能的系统至关重要。尽管多模态大语言模型(MLLMs)展现出强大的视频理解能力,但它们往往更关注语义内容,忽略情感刺激。因此,现有MLLMs在估计观众情绪反应及提供合理解释方面表现不足。为此,我们提出StimuVAR——一种基于MLLM的时空刺激感知视频情感推理框架。StimuVAR采用两级刺激感知机制:帧级感知通过采样最可能引发情绪的视频帧;令牌级感知在令牌空间中进行管状选择,使模型聚焦于情绪触发的时空区域。此外,我们构建了用于情感训练的VAR指令数据,引导MLLM将推理重点转向情感维度,从而增强其情感推理能力。为全面评估有效性,我们设计了包含多种指标的综合评估协议。StimuVAR是首个面向观众中心的情感推理的MLLM方法。实验表明,该方法在理解观众情感反应及生成连贯、深刻的解释方面均具优势。代码已公开于https://github.com/EthanG97/StimuVAR。

原文摘要 · Abstract (English)

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilities, they tend to focus more on the semantic content of videos, often overlooking emotional stimuli. Hence, most existing MLLMs fall short in estimating viewers' emotional reactions and providing plausible explanations. To address this issue, we propose StimuVAR, a spatiotemporal Stimuli-aware framework for Video Affective Reasoning (VAR) with MLLMs. StimuVAR incorporates a two-level stimuli-aware mechanism: frame-level awareness and token-level awareness. Frame-level awareness involves sampling video frames with events that are most likely to evoke viewers' emotions. Token-level awareness performs tube selection in the token space to make the MLLM concentrate on emotion-triggered spatiotemporal regions. Furthermore, we create VAR instruction data to perform affective training, steering MLLMs' reasoning strengths towards emotional focus and thereby enhancing their affective reasoning ability. To thoroughly assess the effectiveness of VAR, we provide a comprehensive evaluation protocol with extensive metrics. StimuVAR is the first MLLM-based method for viewer-centered VAR. Experiments demonstrate its superiority in understanding viewers' emotional responses to videos and providing coherent and insightful explanations. Our code is available at https://github.com/EthanG97/StimuVAR

视频情感大模型多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。