让AI视频生成更讲事实,能解释操作步骤。
Knowledge-Intensive Video Generation

- 用信息查询式提示词驱动视频生成,强调事实性与实用性。
- 构建1080个评测用提示词,自动评估视频事实性与有用性。
- 发现现有模型在流程展示和信息清晰度上远不如人类。
文本到视频生成在视觉质量上进步迅速,但在事实性和实际可用性方面仍缺乏评估。我们提出知识密集型视频生成(KIVI),要求模型根据简短的信息查询类提示生成解释、流程或演示类视频。为此,我们构建了包含1,080个提示的KIVI-Bench基准,并提出用于评估事实性和有用性的自动指标。人工评估显示,我们的指标与人类标注的契合度显著优于现有方法。对七种顶尖视频生成模型的实验表明,当前系统在视觉属性、操作流程和信息呈现清晰度上仍明显落后于人类表现。这些结果凸显了KIVI作为一项具有挑战性的、面向真实场景的视频生成方向。
原文摘要 · Abstract (English)
Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness. We introduce knowledge-intensive video generation (KIVI), where models generate videos from short information-seeking prompts that ask for explanations, procedures, or demonstrations. To evaluate this setting, we construct KIVI-Bench, a benchmark of 1,080 prompts, and propose automatic metrics for factuality and helpfulness. Human evaluation shows that our metrics significantly better align with human annotations than existing alternatives. Experiments on seven state-of-the-art video generation models show that current systems still lag behind human performance, especially on visual properties, procedural operations, and clear information presentation. These results highlight KIVI as a challenging direction for factual and instructionally useful video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。