提出可自适应捕捉多情绪的视频字幕生成框架,提升情感表达准确性。
Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning

- 构建软性情感异构图,融合类别与词汇级情绪节点
- 引入连续门控机制,避免错误分类导致情绪信息丢失
- 联合建模语言生成与情绪分布,支持混合情绪表达
情感视频字幕(EVC)旨在生成既事实准确又具情感表现力的字幕。模型需感知微妙、模糊且随时间变化的情感线索,并将其自然表达为语言,同时不削弱视觉内容。现有方法逐步引入上下文注意力、情感解析、情感先验、动态情感感知与情感因果推理,但大多依赖全局情感向量或刚性层级先验。近期方法采用树状情感先验建立心理情感类别与日常情感词间的粗粒度到细粒度关联,但其硬性子类掩码可能导致粗分类错误时正确词汇情绪被不可逆抑制,且难以表示真实视频中常见的混合或重叠情绪。为此,我们提出SAGML:一种基于情感异构图与多任务语言建模的自适应EVC框架。不再将情感先验视为离散树结构,而是构建包含目录级情感节点与词汇级情感词节点的软性情感异构图。在视频到情感图注意力中注入软门控作为连续偏置,使视觉支持的词汇情感仍可恢复,而非被硬掩码移除。所得情感表征与视觉标记一同输入因果语言解码器,双头(目录与词汇)对提示隐藏状态施加显式情感分布学习。整体模型通过结合自回归字幕生成与情感分布监督的联合目标进行训练。SAGML提供了一种误差鲁棒且多情绪感知的EVC基准。
原文摘要 · Abstract (English)
Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive subtle, ambiguous, and temporally varying emotional cues and translate them into natural language without weakening objective visual content. Existing methods have progressively introduced contextual attention, emotion interpretation, emotion priors, dynamic emotion perception and emotion-cause reasoning. Nevertheless, most of them still depend on either global emotion vectors or rigid hierarchical priors. In recent methods, the tree-structured emotion prior establishes a coarse-to-fine connection between psychological emotion categories and daily emotion words, but its hard subordinate masking may irreversibly suppress correct lexical emotions once the coarse category prediction is inaccurate. It is also limited in representing mixed or overlapping emotions that frequently occur in real videos. To address the issues, we propose SAGML, an adaptive EVC framework via affective heterogeneous graph and multi-task language modeling. Instead of treating the emotion prior as a discrete tree, SAGML constructs a soft affective heterogeneous graph containing catalog-level emotion nodes and lexical-level emotion word nodes. The soft gate is injected into video-to-emotion graph attention as a continuous bias, allowing visually supported lexical emotions to remain recoverable rather than being removed by a hard mask. The resulting affective representation is fed together with visual tokens into a causal language decoder, while dual catalog and lexical heads impose explicit emotion distribution learning on the prompt hidden states. The overall model is trained with a joint objective that combines autoregressive caption generation and emotion distribution supervision. SAGML provides an error-resilient and multi-emotion-aware baseline for EVC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。