arXiv:2503.12077cs.CVcs.AI2025-03CVPR被引 9

用多智能体协作与反思机制,实现复杂视频的开放式风格化。

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

  • 分三步:解析视频分镜、匹配用户描述的风格、多轮自检渲染。
  • 在新基准上超越现有方法6.05%以上,处理复杂转场更稳定。
  • 适合需要自动风格化复杂视频的创作者和研究者。

尽管视频风格化技术取得进展,多数方法仍难以基于开放风格描述处理具有复杂转场的视频。为此,本文提出通用多智能体系统V-Stylist,通过多模态大模型的协作与反思范式实现视频风格化。系统包含三个关键角色:(1) 视频解析器将输入视频分解为多个镜头,并生成关键内容文本提示,采用简洁的视频到镜头提示范式,有效处理复杂转场;(2) 风格解析器识别用户查询中的风格,并通过风格树逐步搜索匹配的风格模型,利用稳健的思维树搜索范式,精确捕捉开放查询中的模糊风格偏好;(3) 风格艺术家使用匹配模型对所有镜头进行风格化渲染,并通过新颖的多轮自反思范式,根据风格要求自适应调整细节控制。该设计模拟人类专业流程,在挑战性任务中取得突破。此外,我们构建了新基准Text-driven Video Stylization Benchmark (TVSBench),填补评估复杂视频风格化在开放查询下的空白。大量实验表明,V-Stylist在整体平均指标上优于FRESCO和ControlVideo,分别提升6.05%和4.51%,标志着视频风格化的重要进展。

原文摘要 · Abstract (English)

Despite the recent advancement in video stylization, most existing methods struggle to render any video with complex transitions, based on an open style description of user query. To fill this gap, we introduce a generic multi-agent system for video stylization, V-Stylist, by a novel collaboration and reflection paradigm of multi-modal large language models. Specifically, our V-Stylist is a systematical workflow with three key roles: (1) Video Parser decomposes the input video into a number of shots and generates their text prompts of key shot content. Via a concise video-to-shot prompting paradigm, it allows our V-Stylist to effectively handle videos with complex transitions. (2) Style Parser identifies the style in the user query and progressively search the matched style model from a style tree. Via a robust tree-of-thought searching paradigm, it allows our V-Stylist to precisely specify vague style preference in the open user query. (3) Style Artist leverages the matched model to render all the video shots into the required style. Via a novel multi-round self-reflection paradigm, it allows our V-Stylist to adaptively adjust detail control, according to the style requirement. With such a distinct design of mimicking human professionals, our V-Stylist achieves a major breakthrough over the primary challenges for effective and automatic video stylization. Moreover,we further construct a new benchmark Text-driven Video Stylization Benchmark (TVSBench), which fills the gap to assess stylization of complex videos on open user queries. Extensive experiments show that, V-Stylist achieves the state-of-the-art, e.g.,V-Stylist surpasses FRESCO and ControlVideo by 6.05% and 4.51% respectively in overall average metrics, marking a significant advance in video stylization.

视频风格化多智能体大模型自反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。