arXiv:2604.10456cs.CV2026-04被引 5

首个电影视频编排基准+多智能体系统,让AI按指令生成连贯短片。

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

论文配图:A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
图 1 · 摘自论文原文
  • 用多智能体重构编排流程,分步设计并组合剧情。
  • 在专业编辑标注的高质量数据上,叙事连贯性显著提升。
  • 适合影视自动化、内容创作与智能编剧研究者。

随着将长篇电影内容改编为短视频的需求激增,亟需通用的自动视频编排系统。然而现有方法局限于预设任务,且社区缺乏全面的评估基准。为此,我们提出 CineBench,首个面向指令驱动电影视频编排的基准,包含多样用户指令和由专业编辑标注的高质量编排结果。为解决上下文坍塌与时间碎片化问题,我们设计 CineAgents,一个将编排重构为“设计-组合”范式的多智能体系统。CineAgents通过剧本逆向工程构建层次化叙事记忆,提供多层次上下文,并采用迭代叙事规划过程,将创意蓝图逐步优化为最终编排脚本。大量实验表明,CineAgents显著优于现有方法,在叙事连贯性和逻辑连贯性上均表现更优。

原文摘要 · Abstract (English)

The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems. However, existing compilation methods are limited to predefined tasks, and the community lacks a comprehensive benchmark to evaluate the cinematic compilation. To address this, we introduce CineBench, the first benchmark for instruction-driven cinematic video compilation, featuring diverse user instructions and high-quality ground-truth compilations annotated by professional editors. To overcome contextual collapse and temporal fragmentation, we present CineAgents, a multi-agent system that reformulates cinematic video compilation into ``design-and-compose'' paradigm. CineAgents performs script reverse-engineering to construct a hierarchical narrative memory to provide multi-level context and employs an iterative narrative planning process that refines a creative blueprint into a final compiled script. Extensive experiments demonstrate that CineAgents significantly outperforms existing methods, generating compilations with superior narrative coherence and logical coherence.

视频编排多智能体电影剪辑指令生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。