首个评估视频编辑理解与操作推理的综合基准,揭示大模型在真实场景中的能力短板。
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

- 构建三轮人机协作标注流程,生成3900段高质量剪辑视频与3080个验证问题对
- 涵盖7种剪辑技巧识别与多片段选择定位任务,实测模型性能远低于人类水平
- 适合视频生成、多模态推理与智能剪辑系统研究者使用
真实世界视频编辑不仅需要电影技法知识,还需多模态推理来选择、对齐并组合画面形成连贯叙事。尽管近期大型多模态模型(LMMs)在通用视频理解上取得显著进展,其在多视频推理与实际编辑工作流中的能力仍鲜有探索。我们提出VEBench,首个专为评估真实视频编辑场景中编辑认知理解与操作推理能力而设计的综合性基准。VEBench包含3.9K段高质量剪辑视频(总时长超257小时)和3,080个人工验证的问答对,通过三轮人机协同标注流程确保时间标签精确与语义一致。该基准包含两项互补任务:1)视频剪辑技巧识别,评估模型利用多模态线索识别7类剪辑技术的能力;2)视频剪辑操作模拟,通过从多个候选片段中选择并精确定位相关片段,模拟真实编辑流程。在专有模型(如Gemini-2.5-Pro)与开源LMM上的广泛实验显示,当前模型表现与人类编辑认知水平存在巨大差距。结果凸显了将视频理解与创造性操作推理相融合的紧迫需求。我们期望VEBench能成为推动智能视频编辑系统发展的基石,并引领未来复杂推理研究。
原文摘要 · Abstract (English)
Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown remarkable progress in general video understanding, their abilities in multi-video reasoning and operational editing workflows remain largely unexplored. We introduce VEBENCH, the first comprehensive benchmark designed to evaluate both editing knowledge understanding and operational reasoning in realistic video editing scenarios. VEBENCH contains 3.9K high-quality edited videos (over 257 hours) and 3,080 human-verified QA pairs, built through a three-round human-AI collaborative annotation pipeline that ensures precise temporal labeling and semantic consistency. It features two complementary QA tasks: 1) Video Editing Technique Recognition, assessing models' ability to identify 7 editing techniques using multimodal cues; and 2) Video Editing Operation Simulation, modeling real-world editing workflows by requiring the selection and temporal localization of relevant clips from multiple candidates. Extensive experiments across proprietary (e.g., Gemini-2.5-Pro) and open-source LMMs reveal a large gap between current model performance and human-level editing cognition. These results highlight the urgent need for bridging video understanding with creative operational reasoning. We envision VEBENCH as a foundation for advancing intelligent video editing systems and driving future research on complex reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。