arXiv:2509.10761cs.CV2025-09International Conf…被引 16

用多智能体自动完成视频非线性剪辑,让语言指令直接生成符合要求的成片。

EditDuet: A Multi-Agent System for Video Non-Linear Editing

  • 设计编辑与评判双智能体,通过语言指令和工具操作实现自动化剪辑。
  • 在用户研究中显著优于现有方法,覆盖更全、时间约束满足率更高。
  • 适合影视制作、短视频创作等需快速产出高质量视频的场景。

视频剪辑自动化在电影制作、广告和社交媒体内容创作中有广泛应用。以往工作主要聚焦于视频检索或用户界面,实际编辑仍依赖人工。本文提出自动化核心剪辑任务,将其建模为序列决策过程。采用多智能体架构:编辑智能体接收视频片段和自然语言指令,使用常见剪辑工具生成序列;评判智能体根据输出提供语言反馈或确认结果。引入基于学习的跨智能体通信机制,实现语言驱动的视频编辑。最后,提出以大模型作为裁判的评估指标,并与人类偏好对比。通过定性和定量用户研究发现,本系统在覆盖度、时间约束满足率及人类偏好上均显著优于现有方法。

原文摘要 · Abstract (English)

Automated tools for video editing and assembly have applications ranging from filmmaking and advertisement to content creation for social media. Previous video editing work has mainly focused on either retrieval or user interfaces, leaving actual editing to the user. In contrast, we propose to automate the core task of video editing, formulating it as sequential decision making process. Ours is a multi-agent approach. We design an Editor agent and a Critic agent. The Editor takes as input a collection of video clips together with natural language instructions and uses tools commonly found in video editing software to produce an edited sequence. On the other hand, the Critic gives natural language feedback to the editor based on the produced sequence or renders it if it is satisfactory. We introduce a learning-based approach for enabling effective communication across specialized agents to address the language-driven video editing task. Finally, we explore an LLM-as-a-judge metric for evaluating the quality of video editing system and compare it with general human preference. We evaluate our system's output video sequences qualitatively and quantitatively through a user study and find that our system vastly outperforms existing approaches in terms of coverage, time constraint satisfaction, and human preference.

视频剪辑多智能体语言驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。