构建首个通用视频编辑基准,支持多维度质量评估。
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

- 提出跨9类32子类的5049例人工标注视频编辑数据集。
- 设计专用奖励模型VEFX-Reward,精准评估指令遵循、渲染质量和编辑局部性。
- 发布300个精选视频提示对,支持主流视频编辑系统标准化对比。
随着AI辅助视频创作日益实用,基于指令的视频编辑成为满足专业需求的关键。然而,该领域仍缺乏大规模人工标注数据集与标准化评估工具。现有资源受限于规模小、缺少编辑输出或缺乏人类质量标签,且评估常依赖昂贵的人工评测或非专业的视觉语言模型。本文提出VEFX-Dataset,一个包含5,049个视频编辑示例的数据集,覆盖9大类32子类,每个样本在指令遵循、渲染质量、编辑局部性三个解耦维度上进行标注。基于此,我们设计了专用于视频编辑质量评估的VEFX-Reward奖励模型,联合处理源视频、编辑指令与编辑后视频,通过序数回归预测各维度得分。进一步发布VEFX-Bench,包含300个精选视频-提示对,实现编辑系统的标准化比较。实验表明,VEFX-Reward在标准IQA/VQA指标及群体偏好评估中均优于通用VLM和已有奖励模型。使用该模型评测主流商业与开源系统,揭示当前模型在视觉合理性、指令遵循与编辑局部性之间存在持续差距。
原文摘要 · Abstract (English)
As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale human-annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision-language model judges that are not specialized for editing quality. We introduce VEFX-Dataset, a human-annotated dataset containing 5,049 video editing examples across 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on VEFX-Dataset, we propose VEFX-Reward, a reward model designed specifically for video editing quality assessment. VEFX-Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores via ordinal regression. We further release VEFX-Bench, a benchmark of 300 curated video-prompt pairs for standardized comparison of editing systems. Experiments show that VEFX-Reward aligns more strongly with human judgments than generic VLM judges and prior reward models on both standard IQA/VQA metrics and group-wise preference evaluation. Using VEFX-Reward as an evaluator, we benchmark representative commercial and open-source video editing systems, revealing a persistent gap between visual plausibility, instruction following, and edit locality in current models. Our project page is https://xiangbogaobarry.github.io/VEFX-Bench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。