新基准测试评估视频模型讲连续故事的能力,发现现有模型表现普遍不佳。
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
- 用7类共423个短故事提示测试模型讲连续事件的能力。
- 11个主流模型平均完成率不足50%,暴露出叙事连贯性短板。
- 适合研究长视频生成、故事理解与多事件一致性问题的学者使用。
当前最先进的视频生成模型虽能产出高细节真实感的视频,但在按提示生成多个连续事件组成的故事情节方面仍表现不佳,而这正是未来长视频生成的关键能力。例如,顶级文本到视频(T2V)模型仍无法准确生成‘如何把大象放进冰箱’这类简单故事。现有以细节为导向的评测主要关注美学质量与时空一致性,却忽略了对事件级故事呈现能力的评估。为此,我们提出StoryEval——一个面向故事完成度的评测基准,包含423个涵盖7类短故事的提示,每条故事由2至4个连续事件构成。我们利用GPT-4V和LLaVA-OV-Chat-72B等先进视觉语言模型,通过一致投票机制验证每个事件是否正确生成,确保评测结果与人工评价高度一致。对11个模型的评估显示,其平均故事完成率未超过50%,揭示了该任务的挑战性。StoryEval为推动T2V模型发展提供了新标准,也指明了下一代连贯故事驱动视频生成的技术方向。
原文摘要 · Abstract (English)
The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts, which is foreseeable an essential capability for future long video generation scenarios. For example, top T2V generative models still fail to generate a video of the short simple story 'how to put an elephant into a refrigerator.' While existing detail-oriented benchmarks primarily focus on fine-grained metrics like aesthetic quality and spatial-temporal consistency, they fall short of evaluating models' abilities to handle event-level story presentation. To address this gap, we introduce StoryEval, a story-oriented benchmark specifically designed to assess text-to-video (T2V) models' story-completion capabilities. StoryEval features 423 prompts spanning 7 classes, each representing short stories composed of 2-4 consecutive events. We employ advanced vision-language models, such as GPT-4V and LLaVA-OV-Chat-72B, to verify the completion of each event in the generated videos, applying a unanimous voting method to enhance reliability. Our methods ensure high alignment with human evaluations, and the evaluation of 11 models reveals its challenge, with none exceeding an average story-completion rate of 50%. StoryEval provides a new benchmark for advancing T2V models and highlights the challenges and opportunities in developing next-generation solutions for coherent story-driven video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。