构建首个覆盖多维度的视频编辑评估基准,解决指令忠实度难题。
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

- 分解视频编辑为时空音参考四大维度,区分显隐指令
- 引入准确率感知惩罚机制,杜绝错误编辑误判高分
- 适合研究视频生成与可控编辑的学者及开发者
基于指令的视频编辑(IVE)是新兴领域,应用广泛,但模型评估仍具挑战。现有基准存在两大缺陷:任务覆盖有限,继承自图像编辑,忽略视频特有维度;评价指标不足,无法有效衡量指令忠实度,导致视觉上合理但实际错误的编辑获得高分。为此,我们提出一个全面、结构化的IVE评估基准。该基准将编辑任务细分为空间、时间、音频和参考式编辑等视频特有维度,超越传统帧级评估,并区分显性与隐性指令,引入基于推理的场景以更贴近真实需求。同时,提出四维评估框架:准确性、保留性、真实性和一致性,结合人工判断与先进视觉语言模型。为强调指令忠实度,设计准确率感知惩罚机制,使其他评分依赖于准确性,防止因原始视频强视觉先验导致错误编辑被高估。在代表性开源与商业模型上的实验表明,当前IVE模型表现仍不理想。OmniEdit-Bench为指令式视频编辑提供了可靠测试平台,并揭示未来研究方向。
原文摘要 · Abstract (English)
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major limitations: limited task coverage inherited from image editing, which overlooks video-specific dimensions, and inadequate metrics that fail to measure instruction fidelity, allowing incorrect edits to receive high scores due to strong visual priors from the original video. To address these issues, we introduce a comprehensive and structured benchmark for IVE. Our benchmark decomposes editing tasks into multiple video-specific dimensions, including spatial, temporal, audio, and reference-based editing, extending beyond conventional frame-level evaluation. It also distinguishes explicit and implicit instructions and incorporates reasoning-based scenarios to better reflect real-world requirements. Furthermore, we propose an evaluation framework that assesses editing quality from four complementary dimensions: accuracy, preservation, realism, and consistency, using both human judgments and state-of-the-art vision-language models. To emphasize instruction fidelity, we introduce an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving inflated evaluations. Extensive experiments on representative open-source and commercial models show that current IVE models remain far from satisfactory. OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。