arXiv:2503.14350cs.CVcs.AI2025-03ICCV被引 40

用指令统一编辑视频,能增删改且理解上下文。

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

论文配图:VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
图 1 · 摘自论文原文
  • 通过多模态大模型理解指令并定位视频内容,生成像素级编辑计划。
  • 在多种编辑任务上超越现有基线,零样本下也能处理新指令。
  • 适合需要灵活视频编辑、跨任务推理的研究者与开发者。

近期视频扩散模型提升了视频编辑能力,但仍难以在统一框架下处理指令式编辑与多样化任务(如添加、移除、更改)。本文提出VEGGIE——基于指令的视频生成式编辑框架,通过端到端设计统一视频概念编辑、语义定位与推理。给定视频和文本指令后,先由多模态大模型解析用户意图,并将其映射到视频上下文,生成每帧对应的编辑查询;随后扩散模型依据这些计划生成符合意图的编辑视频。为支持复杂指令与多任务,采用课程学习策略:先在大规模图像指令数据上对齐多模态大模型与扩散模型,再在高质量多任务视频数据上进行端到端微调。此外,提出一种新型数据合成流程,利用图像到视频模型将静态图像转化为动态、多样化的视频编辑样本。VEGGIE在不同编辑技能上表现优异,优于最佳指令基线,而其他模型在多任务中表现不佳。其在视频对象定位与推理分割方面也显著领先于基线。进一步分析表明,多任务间存在协同效应,展现出零样本多模态指令编辑与上下文内编辑等前景应用。

原文摘要 · Abstract (English)

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we introduce VEGGIE, a Video Editor with Grounded Generation from Instructions, a simple end-to-end framework that unifies video concept editing, grounding, and reasoning based on diverse user instructions. Specifically, given a video and text query, VEGGIE first utilizes an MLLM to interpret user intentions in instructions and ground them to the video contexts, generating frame-specific grounded task queries for pixel-space responses. A diffusion model then renders these plans and generates edited videos that align with user intent. To support diverse tasks and complex instructions, we employ a curriculum learning strategy: first aligning the MLLM and video diffusion model with large-scale instructional image editing data, followed by end-to-end fine-tuning on high-quality multitask video data. Additionally, we introduce a novel data synthesis pipeline to generate paired instructional video editing data for model training. It transforms static image data into diverse, high-quality video editing samples by leveraging Image-to-Video models to inject dynamics. VEGGIE shows strong performance in instructional video editing with different editing skills, outperforming the best instructional baseline as a versatile model, while other models struggle with multi-tasking. VEGGIE also excels in video object grounding and reasoning segmentation, where other baselines fail. We further reveal how the multiple tasks help each other and highlight promising applications like zero-shot multimodal instructional and in-context video editing.

视频编辑指令理解扩散模型多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。