用合成数据训练视频编辑模型,让指令控制视频修改更精准。
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- 用图像编辑器+视频生成器融合生成多样化视频指令数据
- 训练出可执行百万级视频编辑的模型,性能达新纪录
- 适合想做视频智能编辑或数据生成的研究者
基于指令的视频编辑有望降低内容创作门槛,但受限于高质量训练数据稀缺。我们提出Ditto框架,通过融合领先图像编辑器与上下文视频生成器,构建新颖的数据生成流程,突破现有模型覆盖范围局限。为解决成本与质量的矛盾,框架采用轻量级蒸馏模型搭配时序增强模块,显著降低计算开销并提升时间连贯性。整个流程由智能代理驱动,自动设计多样指令并严格筛选输出,实现高质量规模化生成。基于此,我们投入超过12,000 GPU天构建了包含一百万条高保真视频编辑样本的Ditto-1M数据集,并使用课程学习策略在该数据集上训练了Editto模型。实验表明,该模型具备优异的指令遵循能力,在指令式视频编辑任务中达到新的最佳性能。
原文摘要 · Abstract (English)
Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。