用合成参考图提升视频编辑精度,实现精准可控的指令式编辑。
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
- 通过生成模型构建高质量参考图像,扩展训练数据
- 在RefVIE数据集上实现91.2%的指令遵循率和86.5%参考保真度
- 适合需要精细视觉控制的视频编辑研究与应用
基于指令的视频编辑进展迅速,但自然语言难以精确描述复杂视觉细节。参考引导编辑虽具潜力,却受限于高质量配对数据稀缺。为此,我们提出可扩展的数据生成流程,利用图像生成模型将现有视频编辑对转换为高保真训练四元组,构建了面向指令-参考跟随任务的大型数据集RefVIE,并建立RefVIE-Bench用于全面评估。同时,提出统一编辑架构Kiwi-Edit,结合可学习查询与潜在视觉特征实现参考语义引导。通过渐进式多阶段训练,模型在指令遵循与参考保真度上显著提升。大量实验表明,该数据与架构达成了可控视频编辑新基准。所有数据集、模型及代码已开源。
原文摘要 · Abstract (English)
Instruction-based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although reference-guided editing offers a robust solution, its potential is currently bottlenecked by the scarcity of high-quality paired training data. To bridge this gap, we introduce a scalable data generation pipeline that transforms existing video editing pairs into high-fidelity training quadruplets, leveraging image generative models to create synthesized reference scaffolds. Using this pipeline, we construct RefVIE, a large-scale dataset tailored for instruction-reference-following tasks, and establish RefVIE-Bench for comprehensive evaluation. Furthermore, we propose a unified editing architecture, Kiwi-Edit, that synergizes learnable queries and latent visual features for reference semantic guidance. Our model achieves significant gains in instruction following and reference fidelity via a progressive multi-stage training curriculum. Extensive experiments demonstrate that our data and architecture establish a new state-of-the-art in controllable video editing. All datasets, models, and code is released at https://github.com/showlab/Kiwi-Edit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。