arXiv:2607.14202cs.CV2026-07被引 1

首个全面评估关键帧视频生成的基准,揭示模型在忠实还原与画质间的权衡。

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

论文配图:KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
图 1 · 摘自论文原文
  • 构建包含386个样本的多场景基准,覆盖多种条件设置
  • 提出六维关键帧执行指标,结合大模型与感知模型评估质量
  • 发现多数模型在密集关键帧下表现下降,开源模型难理解时间顺序

视频生成越来越多依赖关键帧工作流,创作者通过参考图像序列引导生成过程。尽管近期模型支持多关键帧条件输入,但其能否准确再现指定关键帧并保持整体视频质量仍不明确。本文提出KeyFrame-Compass,首个针对关键帧条件视频生成的综合性评测基准。该基准包含386个精心设计的样本,涵盖三个应用领域、两种视频结构、两种提示粒度、两种条件格式和四种关键帧密度,支持在多样生成环境下进行受控分析。我们进一步提出自动化评估框架,联合衡量关键帧执行与整体视频质量。具体而言,将关键帧执行分解为六个互补指标:存在性、保真度、时序顺序、定位精度、持续性与唯一性;同时通过基于证据的大语言模型判断,并结合专用感知模型评估整体质量。对九个代表性视频生成系统的实验揭示若干根本性局限:当前模型在忠实还原关键帧与自然合成之间存在明显权衡,且随着关键帧约束密度增加,性能显著下降,大多数开源模型也无法将故事板网格输入正确理解为时序关键帧序列。

原文摘要 · Abstract (English)

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.

视频生成关键帧评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。