arXiv:2607.06875cs.CVcs.LG2026-07

构建视频到观众反应分布的映射数据集,助力内容推荐与情感分析。

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

论文配图:Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
图 1 · 摘自论文原文
  • 基于社交媒体构建跨1万+视频的观众情绪分布数据集
  • 利用开源大模型实现86%准确率的低成本持续标注
  • 首次建立真实场景下视频-反应预测基准,适合媒体分析与推荐系统研究

理解并预测观众对视频内容的反应对于提升内容创作、推荐系统和媒体分析至关重要。为此,我们提出《Video2Reaction》,一个跨模态数据集,将短视频片段与真实世界中观众通过社交媒体表达的“诱发情绪”分布进行映射。该数据集包含超过10,000个视频,可作为观众反应预测的可靠基准与训练资源。为实现低成本持续标注(因反应可能随时间变化),我们设计了一种仅使用开源大语言模型的两阶段多智能体流水线,在盲测人类验证下达到86%的正确率,尽管任务本身具有噪声大、主观性强的特点。我们建立了首个真实场景下视频到反应分布预测的基准。实验表明,预训练视频基础模型在零样本设置下表现不佳,但微调后可转化为顶尖预测器,能同时建模完整反应分布与主导反应。然而任务仍具挑战性:最强方法在主导反应预测上仅达77% Top-3 F1(LLaVA-Next),暴露出对集体观众反应建模的显著差距。

原文摘要 · Abstract (English)

Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reaction prediction and other content engagement applications, we introduce $\textbf{Video2Reaction}$, a multimodal dataset that maps short movie segments to a distribution of $\textit{induced emotions}$ of viewers in the wild, as expressed through social media. $\textbf{Video2Reaction}$ spans more than 10,000 videos and serves as a reliable benchmark as well as a training resource for audience reaction prediction. To enable cost-effective continuous annotations as reactions may change over time, we develop a two-stage multi-agent pipeline using only open-source LLMs, achieving 86% correctness under blind human verification despite the inherently noisy and subjective nature of the task. We establish the first benchmark for video-to-reaction-distribution prediction in the wild and show that pretrained foundation video models fail in zero-shot settings, while finetuning transforms them into state-of-the-art predictors capable of modeling both full reaction distributions and dominant responses from video alone. However, the task remains challenging: even the strongest methods achieve only 77% Top-3 F1 in dominant reaction prediction (LLaVA-Next), highlighting a substantial gap in modeling collective audience reaction. \modification{Dataset and code are available at our project page: https://information-fusion-lab-umass.github.io/video2reaction-bench.github.io

视频生成情感分析多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。