通过细粒度情感-原因配对提取,提升视频描述的情感准确性。
Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

- 分两阶段提取视觉与情感特征,增强语义理解
- 在EVC-MSVD上比基线提升4.4% BLEU-2和5.4% ROUGE-L
- 适合关注情感感知与视频生成的开发者
情感视频描述(EVC)旨在生成事实准确且情感丰富的视频描述。现有方法依赖全局视觉特征捕捉整体情感线索,再融合多模态特征引导生成,但忽略了情感由特定动机引发的关键特性——这些动机通常仅存在于核心视频片段中。全局提取导致信息冗余和情感线索失准。为此,本文提出一种细粒度情感-原因配对提取框架。首先,设计概念感知的视觉语义分解模块,通过场景、物体和运动概念增强视觉特征;同时提出视觉引导的情感可解释学习模块,利用视觉时序动态优化情感表示,并引入可靠的VAD向量约束增强可解释性。其次,通过前后细化特征的交叉耦合实现情感-原因配对提取,并采用对比损失实现语义强制对齐。在三个挑战性数据集上的实验表明,该方法显著优于基线,尤其在EVC-MSVD上分别取得+4.4%(BLEU-2)和+5.4%(ROUGE-L)的提升,验证了各模块的有效性。
原文摘要 · Abstract (English)
Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotional cues, and then aggregate multimodal features to guide the emotional caption generation, which ignores the critical characteristic of the EVC task. Visual emotions are evoked by specific motivational causes, which are usually only implied in core video segments. The holistic mining brings significant information redundancy and inaccurate emotional cues. Thus, fine-grained visual cause extraction has a facilitative effect on both emotion perception and emotion-attributed caption generation. To this end, we propose a fine-grained emotion-cause pair extraction framework for emotion-attributed video captioning. Specifically, we learn pair-wise emotion and cause features in two rounds: 1) We propose a Concept-aware Visual Semantic Decomposition module to augment visual features by exploring scene, object, and motion concepts. Besides, to enhance emotional features, we propose a Visual-guided Emotion Interpretable Learning module, which guides emotion refinement with visual temporal dynamics, and augments the interpretable refinement process by reliable VAD-vector constraints. 2) We achieve emotion-cause pair extraction by cross-coupling the visual and emotional features before and after refinement, and leverage contrastive loss to achieve semantic forced alignment. Overall, our approach optimizes complex semantic understanding and emotion perception of videos, leading to a promising performance in emotional captioning. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, e.g., achieving the best performances with +4.4% and +5.4% w.r.t. BLEU-2 and ROUGE-L, respectively, on the EVC-MSVD dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。