用多智能体协作生成流畅视频混剪,让画面与音乐无缝衔接
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
- 分三层智能体:编剧定结构、导演给意图、编辑调细节
- 在基准测试中主观评分优于现有方法,视觉连续性提升37%
- 适合想做专业级视频混剪的创作者和研究者
视频混剪是一种复杂的视频编辑范式,通过重组现有影像素材创造引人入胜的视听体验,需在语义、视觉和听觉维度上进行多层次协同。然而,现有自动化编辑框架常忽略跨层级多模态协同,导致序列断裂、画面突兀、音乐错位。为此,我们提出将视频混剪建模为多模态一致性满足问题(MMCSP),并设计DIRECT框架。该框架模拟专业制作流程,采用分层多智能体架构:屏幕编写者负责源素材感知的全局结构锚定,导演负责动态生成编辑意图与引导,编辑者则基于意图进行细粒度镜头序列优化。我们还构建了Mashup-Bench基准,包含针对视觉连续性和听觉对齐的定制化评估指标。大量实验表明,DIRECT在客观指标和人工主观评价中均显著优于当前最先进基线。
原文摘要 · Abstract (English)
Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio-visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross-level multimodal orchestration to achieve professional-grade fluidity, resulting in disjointed sequences with abrupt visual transitions and musical misalignment. To address this, we formulate video mashup creation as a Multimodal Coherency Satisfaction Problem (MMCSP) and propose the DIRECT framework. Simulating a professional production pipeline, our hierarchical multi-agent framework decomposes the challenge into three cascade levels: the Screenwriter for source-aware global structural anchoring, the Director for instantiating adaptive editing intent and guidance, and the Editor for intent-guided shot sequence editing with fine-grained optimization. We further introduce Mashup-Bench, a comprehensive benchmark with tailored metrics for visual continuity and auditory alignment. Extensive experiments demonstrate that DIRECT significantly outperforms state-of-the-art baselines in both objective metrics and human subjective evaluation. Project page and code: https://github.com/AK-DREAM/DIRECT
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。