arXiv:2604.05076cs.MAcs.MM2026-04被引 1

用多智能体协作实现音乐驱动的复杂视频混剪,自动对齐节奏与叙事。

GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing

  • 分层双循环架构:全局规划+局部精修,协同处理长时序编辑任务。
  • 在两个任务设置上分别超越最强基线33.2%和15.6%,显著提升生成质量。
  • 适合需要高自由度、强节奏对齐的视频创作场景,如短视频混剪、MV制作。

音乐驱动的混剪视频创作是极具挑战性的非线性视频编辑任务,系统需从大量源视频中组合出连贯时间线,同时匹配音乐节奏、用户意图、故事完整性及长程结构约束。现有方法多依赖固定流程或简化检索拼接范式,难以适应多样提示与异构素材。本文提出GLANCE,一种用于音乐驱动非线性视频编辑的全局-局部协调多智能体框架。其采用双循环架构:外环进行长程规划与任务图构建,内环通过“观察-思考-行动-验证”流程完成分段编辑及迭代优化。为解决子时间线合并后的跨段冲突与全局不一致问题,引入兼具预防与纠正功能的全局-局部协调机制,包含新型上下文控制器、冲突区域分解模块及自底向上的动态协商机制。为支持严谨评估,构建了MVEBench新基准,按任务类型、提示明确度与音乐长度分解编辑难度,并提出代理作为裁判的评估框架。实验表明,GLANCE在相同骨干模型下持续优于先前基线与开源产品基线。以GPT-4o-mini为骨干,其在两项任务设置上分别提升33.2%和15.6%。人工评估进一步验证生成视频质量及评估框架有效性。

原文摘要 · Abstract (English)

Music-grounded mashup video creation is a challenging form of video non-linear editing, where a system must compose a coherent timeline from large collections of source videos while aligning with music rhythm, user intent, story completeness, and long-range structural constraints. Existing approaches typically rely on fixed pipelines or simplified retrieval-and-concatenation paradigms, limiting their ability to adapt to diverse prompts and heterogeneous source materials. In this paper, we present GLANCE, a global-local coordination multi-agent framework for music-grounded nonlinear video editing. GLANCE adopts a bi-loop architecture for better editing practice: an outer loop performs long-horizon planning and task-graph construction, and an inner loop adopts the "Observe-Think-Act-Verify" flow for segment-wise editing tasks and their refinements. To address the cross-segment and global conflict emerging after subtimelines composition, we introduce a dedicated global-local coordination mechanism with both preventive and corrective components, which includes a novelly designed context controller, conflict region decomposition module, and a bottom-up dynamic negotiation mechanism. To support rigorous evaluation, we construct MVEBench, a new benchmark that factorizes editing difficulty along task type, prompt specificity, and music length, and propose an agent-as-a-judge evaluation framework for scalable multi-dimensional assessment. Experimental results show that GLANCE consistently outperforms prior research baselines and open-source product baselines under the same backbone models. With GPT-4o-mini as the backbone, GLANCE improves over the strongest baseline by 33.2% and 15.6% on two task settings, respectively. Human evaluation further confirms the quality of the generated videos and validates the effectiveness of the proposed evaluation framework.

视频生成多智能体音乐对齐非线性编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。