arXiv:2603.17455cs.CV2026-03被引 5

解决视频描述中事实与情感失衡问题,提升生成准确性。

FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning

  • 引入外部语料库检索增强语义,统一框架协同挖掘事实与情感。
  • 通过不确定性估计校准事实,动态调整不同样本的偏倚程度。
  • 基于校准事实生成视觉查询,自适应增强情感表达,适合多模态生成研究者。

情感视频字幕(EVC)是一项新兴任务,旨在描述视频中的事实内容并表达其内在情感。现有方法通常感知全局情感线索,并与视频内容结合生成描述,但因事实与情感线索挖掘不足且生成过程协调性差,难以应对事实-情感偏差问题——即不同样本在生成时事实与情感需求存在差异。为此,本文提出一种检索增强的框架FACE-net,通过统一架构协同挖掘事实-情感语义,并为生成提供自适应、精准引导,突破了全样本学习中事实-情感描述的妥协倾向。技术上,首先引入外部语料库,检索与视频内容最相关的句子以增强语义信息;随后,通过不确定性估计模块进行事实校准,将检索信息拆分为主谓宾三元组,并利用视频内容进行自反与交叉精炼,有效挖掘事实语义;同时,渐进式视觉情感增强模块以校准后的事实语义为专家,与视频内容和情感词典交互生成视觉查询与候选情感,再聚合以自适应增强各事实语义的情感表达。此外,为缓解事实-情感偏差,设计动态偏倚调节路由模块,预测并调整样本的偏倚程度。

原文摘要 · Abstract (English)

Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotional cues and then combine with video content to generate descriptions. However, insufficient factual and emotional cues mining and coordination during generation make their methods difficult to deal with the factual-emotional bias, which refers to the factual and emotional requirements being different in different samples on generation. To this end, we propose a retrieval-enhanced framework with FActual Calibration and Emotion augmentation (FACE-net), which through a unified architecture collaboratively mines factual-emotional semantics and provides adaptive and accurate guidance for generation, breaking through the compromising tendency of factual-emotional descriptions in all sample learning. Technically, we firstly introduces an external repository and retrieves the most relevant sentences with the video content to augment the semantic information. Subsequently, our factual calibration via uncertainty estimation module splits the retrieved information into subject-predicate-object triplets, and self-refines and cross-refines different components through video content to effectively mine the factual semantics; while our progressive visual emotion augmentation module leverages the calibrated factual semantics as experts, interacts with the video content and emotion dictionary to generate visual queries and candidate emotions, and then aggregates them to adaptively augment emotions to each factual semantics. Moreover, to alleviate the factual-emotional bias, we design a dynamic bias adjustment routing module to predict and adjust the degree of bias of a sample.

视频生成情感理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。