arXiv:2511.02712cs.CV2025-11NeurIPS被引 10

构建情感树推理框架,提升视频情绪理解的可解释性与准确性。

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

  • 分阶段融合情绪属性感知、表达分析与高层理解,构建情感树推理机制。
  • 在15项人脸情绪任务上达到新基准,210万条指令样本支持细粒度分析。
  • 适合情绪计算、人机交互及视频理解领域研究者参考。

从视频中理解与预测情绪近年来受到广泛关注,得益于视频大语言模型(VideoLLMs)的发展。尽管先进方法在视频情绪分析上取得进展,但情绪固有的动态性和依赖线索的特性仍带来挑战,难以合理解释复杂演变的情绪状态。为此,我们提出一种新型情感线索引导的推理框架,以分阶段方式统一基础属性感知、表情分析与高层情绪理解。核心是专为情绪推理与指令遵循设计的视频情绪基础模型(VidEmo),采用两阶段微调:先进行课程化情绪学习注入情绪知识,再通过情感树强化学习实现情绪推理。同时,我们建立基础数据基础设施,引入以情绪为中心的细粒度数据集(Emo-CFG),包含210万条多样化指令样本,涵盖可解释的情绪问答、细粒度描述及对应推理理由,为推进情绪理解任务提供关键资源。实验表明,该方法在15项人脸感知任务上表现优异,树立新里程碑。

原文摘要 · Abstract (English)

Understanding and predicting emotion from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions poses significant challenges. Emotions are characterized by dynamic and cues-dependent properties, making it difficult to understand complex and evolving emotional states with reasonable rationale. To tackle these challenges, we propose a novel affective cues-guided reasoning framework that unifies fundamental attribute perception, expression analysis, and high-level emotional understanding in a stage-wise manner. At the core of our approach is a family of video emotion foundation models (VidEmo), specifically designed for emotion reasoning and instruction-following. These models undergo a two-stage tuning process: first, curriculum emotion learning for injecting emotion knowledge, followed by affective-tree reinforcement learning for emotion reasoning. Moreover, we establish a foundational data infrastructure and introduce a emotion-centric fine-grained dataset (Emo-CFG) consisting of 2.1M diverse instruction-based samples. Emo-CFG includes explainable emotional question-answering, fine-grained captions, and associated rationales, providing essential resources for advancing emotion understanding tasks. Experimental results demonstrate that our approach achieves competitive performance, setting a new milestone across 15 face perception tasks.

情绪识别视频理解推理框架基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。