首个针对长视频叙事能力的综合评估基准,量化叙事丰富度。
NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
- 定义时间叙事原子(TNA)作为叙事单元,量化视频叙事密度。
- 构建自动提示生成管道,支持可扩展的多阶段叙事测试。
- 基于大模型问答框架,评估模型在复杂叙事表达中的表现。
随着基础视频生成技术的快速发展,长视频生成模型展现出广阔的研究潜力,其目标不仅是延长视频时长,更需准确表达更丰富的叙事内容。然而,由于缺乏专门针对长视频生成模型的评估基准,现有评估仍依赖于简单叙事提示的基准(如VBench)。本文提出的NarrLV是首个全面评估长视频生成模型叙事表达能力的基准。受电影叙事理论启发,我们首先引入时间叙事原子(TNA)作为保持视觉连贯性的基本叙事单位,并用其数量定量衡量叙事丰富度;基于影响TNA变化的三个关键电影叙事元素,构建了可灵活扩展叙事层级的自动提示生成管道。其次,依据叙事表达的三个渐进层次,设计了基于多模态大模型的问答式评估指标。最后,在多个现有长视频生成模型与基础生成模型上进行了广泛评估,实验结果表明该指标与人类判断高度一致,揭示了当前模型在叙事表达方面的具体能力边界。
原文摘要 · Abstract (English)
With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video generation tasks is not only to extend video duration but also to accurately express richer narrative content within longer videos. However, due to the lack of evaluation benchmarks specifically designed for long video generation models, the current assessment of these models primarily relies on benchmarks with simple narrative prompts (e.g., VBench). To the best of our knowledge, our proposed NarrLV is the first benchmark to comprehensively evaluate the Narrative expression capabilities of Long Video generation models. Inspired by film narrative theory, (i) we first introduce the basic narrative unit maintaining continuous visual presentation in videos as Temporal Narrative Atom (TNA), and use its count to quantitatively measure narrative richness. Guided by three key film narrative elements influencing TNA changes, we construct an automatic prompt generation pipeline capable of producing evaluation prompts with a flexibly expandable number of TNAs. (ii) Then, based on the three progressive levels of narrative content expression, we design an effective evaluation metric using the MLLM-based question generation and answering framework. (iii) Finally, we conduct extensive evaluations on existing long video generation models and the foundation generation models. Experimental results demonstrate that our metric aligns closely with human judgments. The derived evaluation outcomes reveal the detailed capability boundaries of current video generation models in narrative content expression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。