arXiv:2606.16198cs.CV2026-06

通过显式提取动作事件与视觉证据,提升视频情感预测的精准度。

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

论文配图:GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction
图 1 · 摘自论文原文
  • 从视频描述中提取时序化的主谓宾三元组和文本线索
  • 将人物物体实体定位为视觉裁块,提供具体视觉证据
  • 适用于需要细粒度情感分析的广告与影视视频场景

视频广告中的观众情感预测旨在推断受众隐含的情绪反应。为弥合呈现内容与感知情绪之间的差距,模型需从显性视觉叙事、具体的人物-物体交互及可见文本线索中推断隐藏情绪。然而,标准多模态大模型通常依赖整体帧表示,使这些细粒度的情感相关事件模糊化,阻碍精确情绪推理。为此,我们提出一种基于动作中心的显式证据增强框架,通过引入结构化事件与局部视觉证据,提升视频多模态大模型的线索提取与理解能力。该方法从动作中心的视频描述中提取时序排列的主谓宾(SVO)三元组和辅助可见文本线索,将主体与客体实体定位为视觉实体裁块,并引导多模态大模型基于这些结构化线索进行增强型情绪推理。其中,动作三元组明确‘发生了什么’,而接地的视觉裁块则锚定‘谁或什么参与了每个事件’于具体视觉证据。在Pitts数据集上的实验表明,该方法持续优于Qwen2.5-VL与Qwen3-VL基线。消融实验、在AdsQA上的跨数据集评估以及针对情感聚焦的TVQA子集的迁移实验进一步验证了方法的有效性与泛化能力。

原文摘要 · Abstract (English)

Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, models must deduce hidden viewer emotions from explicit visual narratives, concrete character-object interactions, and visible textual cues. However, standard Multimodal Large Language Models (MLLMs) typically rely on holistic frame representations, which leave these fine-grained, affect-relevant events implicit and complicate precise emotional reasoning. To address this, we propose a grounded action-centric evidence augmentation framework that enhances video MLLMs' clue extraction and comprehension by introducing explicit event structure and localized visual evidence. Our method extracts temporally ordered subject-verb-object (SVO) triplets and auxiliary visible textual cues from action-centric video descriptions, grounds subject and object entities as visual entity crops, and then enables the MLLM to perform clue-enhanced emotional reasoning based on these extracted structured clues. In this way, action triplets specify "what happens", while grounded visual entity crops anchor "who or what participates in each event" to concrete visual evidence. Experiments on the Pitts dataset show consistent improvements over Qwen2.5-VL and Qwen3-VL baselines. Ablation studies, cross-dataset evaluation on AdsQA, and transfer experiments on an emotion-focused TVQA subset further support the effectiveness and generalization of our approach.

视频情感多模态动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。