arXiv:2511.11002cs.CV2025-11中稿 · AAAI被引 2

首个面向创意视频的情感数据集,助力生成更富情感表达的动画与非现实视频。

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

  • 构建包含动画、电影片段等的多模态情感标注视频数据集。
  • 发现视觉特征与情绪感知在时空上的关联模式,提升生成质量。
  • 适合从事情感计算、视频生成与艺术创作的研究者使用。

情感在视频表达中至关重要,但现有视频生成系统多关注低层视觉指标,忽视情感维度。尽管视觉领域的情感分析已有进展,视频社区仍缺乏将情感理解与生成任务结合的专用资源,尤其在风格化和非现实场景中。为此,我们提出 EmoVid,首个专为创意媒体设计的多模态情感标注视频数据集,涵盖卡通动画、电影片段与动态贴纸。每段视频均标注情感标签、视觉属性(亮度、色彩度、色相)及文本描述。通过系统分析,我们揭示了不同视频形式中视觉特征与情绪感知之间的时空模式。基于这些发现,我们通过对 Wan2.1 模型进行微调,构建情感条件下的视频生成方法。结果表明,在文本到视频与图像到视频任务中,生成视频在定量指标与视觉质量上均有显著提升。EmoVid 为情感化视频计算建立了新基准。本工作不仅深化了对艺术风格视频中视觉情感分析的理解,也为增强视频生成中的情感表达提供了实用方法。

原文摘要 · Abstract (English)

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistically styled videos, but also provides practical methods for enhancing emotional expression in video generation.

情感计算视频生成多模态数据集艺术风格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。