新基准CoVEBench测试视频模型处理复杂指令的能力
CoVEBench: Can Video Editing Models Handle Complex Instructions?

- 构建包含626条多点指令的综合性视频编辑评测集
- 现有模型在复合编辑中常漏改、破坏原内容或引入伪影
- 适合研究真实场景下视频编辑模型的开发者使用
尽管近期文本引导的视频编辑模型在基础任务(如风格迁移、物体插入)上表现优异,但现实用户请求高度复合化。单一提示常需同时修改主体、动作和镜头视角,同时严格保留无关时空内容。现有基准受限于单一编辑和粗粒度全局指标,无法诊断模型对复杂流程的处理能力。为此,我们提出CoVEBench,一个包含416个精选源视频、626条多点编辑指令和9,990项细粒度检查项的复合式视频编辑基准。涵盖多样编辑维度,通过MLLM判断指令符合度与视频保真度,并辅以自动化视频质量指标。大量实验表明,复合编辑仍是重大挑战:当前模型在同时执行多项操作时,常遗漏修改、违反保留约束或引入伪影。CoVEBench提供了一个具有挑战性且可诊断的测试平台,推动视频编辑向真实用户工作流发展。
原文摘要 · Abstract (English)
While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt often demands multiple coupled edits, such as modifying subjects, actions, and camera views, while strictly preserving unrelated spatiotemporal content. Existing benchmarks, heavily constrained by isolated edits and coarse global metrics, fail to diagnose how models handle such complex workflows. To address this gap, we introduce CoVEBench, a compositional video editing benchmark comprising 416 curated source videos, 626 multi-point editing instructions, and 9,990 fine-grained checklist items. Covering diverse editing dimensions, CoVEBench evaluates models via MLLM-judged instruction compliance and video fidelity, alongside automated metrics for video quality. Extensive experiments reveal that compositional editing remains a profound challenge: current models frequently omit edits, violate preservation constraints, or introduce artifacts when handling multiple operations simultaneously. CoVEBench provides a challenging, diagnostic testbed to advance video editing toward realistic user workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。