首个评测多镜头音视频生成中剪辑技巧执行能力的基准,揭示当前模型虽流畅但难精准实现专业剪辑。
Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation

- 设计带明确剪辑指令的结构化提示,构建可量化剪辑执行的评估框架
- 13个主流模型普遍存在剪辑指令执行失败,高阶蒙太奇效果下降明显
- 适合研究音视频生成、影视自动化与交互式内容创作的开发者与学者
近期多镜头音视频生成器能产出日益连贯且具有电影感的输出,但连贯性并不等于具备剪辑技巧执行能力。专业剪辑依赖于镜头结构、转场语法、音画剪切关系及蒙太奇手法,而现有基准大多依赖内容质量、同步性或物理合理性等代理指标,系统性忽略了编辑指令是否被实际执行。我们提出CutCraft,首个针对多镜头音视频生成中剪辑技巧执行的基准。CutCraft通过在结构化多镜头提示中加入显式剪辑规范,并配套分层混合评估框架,结合镜头结构对齐、专家模型度量、工具引导的多模态判断及基于评分标准的问题回答。此外,我们设计了一个代理式剪辑基线,将生成分解为规划、镜头级合成与后置组合,显式实现如J-cut、L-cut和转场时机等剪辑语义。在13个先进闭源与开源模型上测试显示,当前系统普遍在连贯性与剪辑执行间存在显著差距:虽能生成合理多镜头视频,但难以可靠执行编辑指令。我们发现镜头结构不稳定、转场控制薄弱,且高阶蒙太奇性能急剧下降,而美学质量与剪辑合规性仅弱相关。该基准、评估方法及剪辑代理基线已开源至https://github.com/AlibabaResearch/cut-craft-bench。
原文摘要 · Abstract (English)
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute editing techniques. Professional editing depends on shot structure, transition grammar, audio-video cut relations, and montage, yet existing benchmarks largely rely on proxies such as content quality, synchronization, or physical plausibility, systematically missing whether such editing instructions are actually executed. We introduce CutCraft, the first benchmark for editing-technique execution in multi-shot audio-video generation. CutCraft extends structured multi-shot prompts with explicit editing specifications and is paired with a hierarchical hybrid evaluation framework that combines shot-structure alignment, expert-model metrics, tool-grounded multimodal judgment, and rubric-based question answering. Beyond evaluation, we design an agentic editing baseline that decomposes generation into planning, shot-level synthesis, and post-hoc composition, explicitly realizing editing semantics such as J-cuts, L-cuts, and transition timing. Across 13 state-of-the-art closed- and open-source models, CutCraft reveals a consistent gap between coherence and editing-technique execution: current systems often produce plausible multi-shot videos yet fail to execute editorial instructions reliably. We find unstable shot structures, weak control of transition execution, and sharp degradation on higher-order montage, while aesthetic quality is only weakly correlated with editing-technique compliance. The benchmark and metrics, and the editing agent baseline are available at https://github.com/AlibabaResearch/cut-craft-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。