arXiv:2506.01689cs.AIcs.CL2025-06被引 2

构建真实用户视频问答基准,评估文本生成视频能力。

Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents

  • 从真实对话中提取7500条需视频回答的查询
  • 建立4500组高质量图文匹配对,覆盖多样化场景
  • 适合研究多模态生成、视频理解与用户意图建模者

通过大语言模型获取信息已成为主流方式,但现有问答数据集主要聚焦文本响应,难以应对需要视觉演示或解释的复杂用户查询。为填补这一空白,我们构建了名为RealVideoQuest的基准,用于评估文本到视频(T2V)模型在回应现实世界、视觉语境化查询方面的能力。该基准从Chatbot-Arena中识别出7500条具有视频响应意图的真实用户查询,并通过多阶段视频检索与精炼流程,构建了4500组高质量的查询-视频配对。我们进一步开发了多角度评估体系,以衡量生成视频答案的质量。实验表明,当前T2V模型在有效回应真实用户查询方面仍存在显著困难,揭示了多模态AI中的关键挑战与未来研究方向。

原文摘要 · Abstract (English)

Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focus on textual responses, making it challenging to address complex user queries that require visual demonstrations or explanations for better understanding. To bridge this gap, we construct a benchmark, RealVideoQuest, designed to evaluate the abilities of text-to-video (T2V) models in answering real-world, visually grounded queries. It identifies 7.5K real user queries with video response intents from Chatbot-Arena and builds 4.5K high-quality query-video pairs through a multistage video retrieval and refinement process. We further develop a multi-angle evaluation system to assess the quality of generated video answers. Experiments indicate that current T2V models struggle with effectively addressing real user queries, pointing to key challenges and future research opportunities in multimodal AI.

视频生成多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。