arXiv:2609.01816cs.CV2026-09

构建视频情绪反应数据集,让模型学会预测观众真实情感反馈。

Video2Reaction: Training Foundation Video Models to Predict Audience Reaction

论文配图:Video2Reaction: Training Foundation Video Models to Predict Audience Reaction
图 1 · 摘自论文原文
  • 用社交媒体评论构建多模态情绪标签数据集
  • 微调后模型在情绪预测上超越专用基线
  • 预训练模型仅用1%数据即可达到顶尖性能

我们提出Video2Reaction,一个将短视频片段与真实观众情绪反应映射的多模态数据集,通过大规模聚合社交媒体评论获取自然的情绪多样性。情绪标签以类别分布形式建模,更准确反映情感感知的主观性和模糊性。我们在两个视觉-语言模型(VLM)上进行LoRA微调,结果表明这些模型能有效学习该数据集,并在主流情绪预测任务中优于专门设计的基线模型。进一步实验显示,经过Video2Reaction预微调的模型可高效迁移至另一个情绪数据集VCE(不同分类体系和视频领域),其中仅使用1%的VCE训练数据,LLaVA-NeXT-Video-7B模型即达到0.682的top-3准确率,与全量数据训练的最佳表现相当。

原文摘要 · Abstract (English)

We introduce Video2Reaction, a multimodal dataset that maps short movie segments to the induced emotional reactions of viewers in the wild, as expressed through social media comments. Video2Reaction captures the natural diversity of emotional responses by aggregating reactions from online comments at scale, modeling labels as distributions over categorical emotions to better reflect the subjective and ambiguous nature of emotional perception. We benchmark two vision-language models (VLMs) finetuned with LoRA, showing that VLMs learn effectively from Video2Reaction and outperform specialized baselines on dominant reaction prediction. We further demonstrate that VLMs pre-finetuned on Video2Reaction transfer effectively to VCE, another induced emotion dataset with a different taxonomy and video domain. Notably, LLaVA-NeXT-Video-7B pre-finetuned on Video2Reaction and adapted on only 1% of VCE training data achieves a top-3 accuracy of 0.682, on par with the best reported VCE performance trained on the full dataset. The dataset is available at https://huggingface.co/datasets/infofusionlab/Video2Reaction

视频情绪多模态迁移学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。