首个面向观众情感变化的视频情绪动态推理基准,支持细粒度情绪追踪与因果分析。
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset

- 构建动态情感推理框架,基于连续事件建模情绪演变过程。
- 包含1.5万段视频与3.7万条情绪片段,覆盖27类情绪及因果链标注。
- 适用于多模态大模型在情感理解、视频叙事分析中的研究与评估。
视频情绪分析通常被当作静态分类任务处理,将每段视频视为独立标注单元。然而,这种设定忽略了心理学核心事实:情绪会随连续因果事件的累积反应而变化。为此,我们提出动态情感推理(Dynamic Affective Reasoning),这是首个大规模、面向观众的视频情绪转变与因果推理基准。DAR包含15,087个视频和36,908个事件对齐的情绪片段,标注了27种情绪类别。不同于现有视频情绪数据集,DAR从观众视角出发,提供细粒度情绪表达与转变的密集标注,具备时间精准性与因果显式性。基于此,我们正式定义三项挑战性任务:情绪分割、细粒度情绪分类与情感推理。为应对该基准,我们提出DAR-R1——一种两阶段框架,结合监督微调与组相对策略优化。在10+多模态大模型上的实验表明,DAR-R1在情绪定位与情感推理上均达到新最佳性能。项目页面:https://github.com/Zhang-Zhiyan/DAR。
原文摘要 · Abstract (English)
Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psychological fact: emotions change as a result of cumulative reactions to consecutive causal events. To bridge this gap, we introduce Dynamic Affective Reasoning, the first large-scale benchmark for viewer-centric affect transitions and causal reasoning over consecutive video events. DAR contains 15,087 videos and 36,908 event-aligned affective segments annotated with 27 emotion categories. Unlike existing video-based emotion datasets, DAR presents a new viewer-centric perspective on fine-grained emotional expressions and transitions, and provides dense, temporally grounded, and causally explicit reasoning chains. Based on DAR, we formally define three challenging tasks: affective segmentation, fine-grained emotion classification, and affective reasoning. Complementing this benchmark, we propose DAR-R1, a two-stage framework that combines supervised fine-tuning with Group Relative Policy Optimization. Experiments across 10+ MLLMs show that DAR-R1 sets a new state-of-the-art for dynamic affective reasoning, in terms of both emotional localization and affective reasoning. Project page: https://github.com/Zhang-Zhiyan/DAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。