构建视频生成评估新基准,精准衡量编辑一致性与视觉质量。
V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

- 设计11维多维度评估框架,覆盖时序对齐等关键指标。
- 在6个专属维度上与人工评价相关性达0.905,接近人类判断。
- 对比三款模型,揭示不同模型在编辑精度与画质上的互补优势。
视频到视频(V2V)生成难以评估,因其输出需同时遵循编辑指令并保持与源视频的帧级对应关系,而现有文本到视频和图像到视频的度量方法无法捕捉这些特性。本文提出 V2V-Bench,一个包含11个维度的综合性评估基准,分为五个类别:时序对齐、结构保真度、变换质量、视频质量和语义对齐。该基准将多样化的源视频与具有挑战性的编辑任务配对,评估两款商用模型(Grok Imagine、Gemini Veo3)和一个开源模型(Open Sora 2)。结果表明,各模型表现互补:Grok 在编辑保真度上更优,Veo3 则在视觉质量上更强。在六个专用于 V2V 的维度上,V2V-Bench 与人工评价的斯皮尔曼相关系数达到 0.905。
原文摘要 · Abstract (English)
Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence with the source video, which existing T2V and I2V metrics do not capture. We introduce V2V-Bench, a 11-dimension benchmark organized into five categories: temporal alignment, structural fidelity, transformation quality, video quality, and semantic alignment. V2V-Bench pairs diverse source videos with challenging editing tasks and evaluates two commercial models, Grok Imagine and Gemini Veo3, and one open-source model, Open Sora 2. Results show complementary model strengths: Grok performs better on editing fidelity, while Veo3 achieves stronger visual quality. On six V2V-specific dimensions, V2V-Bench reaches a Spearman correlation of 0.905 with human judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。