用视觉语言模型指导视频编辑,让AI更懂复杂指令。
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization

- 用VLM融合指令、首帧和参考图,生成精准语义表示。
- 通过相对奖励优化,显著提升指令遵循与画面一致性。
- 适合需要高精度自然语言控制视频编辑的开发者或创作者。
基于指令的视频编辑旨在根据自然语言指令修改输入视频,同时保持内容保真度和时间连贯性。然而,现有基于扩散模型的方法通常在简单编辑操作的成对数据上训练,从根本上限制了其对多样化、复杂真实指令的泛化能力。为解决这一泛化差距,我们提出VIVA,一个可扩展的指令式视频编辑框架,结合视觉语言模型(VLM)引导编码与奖励优化。首先,引入基于VLM的指令器,将文本指令、源视频首帧及可选参考图像编码为视觉接地的指令表征,为扩散变压器主干提供细粒度的空间与语义上下文。其次,提出后训练阶段Edit-GRPO,将组相对策略优化(Group Relative Policy Optimization)适配至视频编辑领域,直接利用相对奖励优化模型,使其输出更忠实于指令、保留内容且具有美学质量。此外,设计了一条数据构建流水线,用于合成多样、高保真的基础编辑操作成对视频-指令数据。大量实验表明,VIVA在指令遵循、泛化能力和编辑质量方面均优于现有最先进方法。
原文摘要 · Abstract (English)
Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired data of simple editing operations, which fundamentally limits their ability to generalize to diverse and complex, real-world instructions. To address this generalization gap, we propose VIVA, a scalable framework for instruction-based video editing that leverages VLM-guided encoding and reward optimization. First, we introduce a VLM-based instructor that encodes the textual instruction, the first frame of the source video, and an optional reference image into visually-grounded instruction representations, providing fine-grained spatial and semantic context for the diffusion transformer backbone. Second, we propose a post-training stage, Edit-GRPO, which adapts Group Relative Policy Optimization to the domain of video editing, directly optimizing the model for instruction-faithful, content-preserving, and aesthetically pleasing edits using relative rewards. Furthermore, we propose a data construction pipeline designed to synthetically generate diverse, high-fidelity paired video-instruction data of basic editing operations. Extensive experiments show that VIVA achieves superior instruction following, generalization, and editing quality over state-of-the-art methods. Website: https://viva-paper.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。