构建200万对高质量视频编辑数据,提升端到端模型效果
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
- 用4个专业级模型生成200万组高质量视频编辑对
- 经筛选后数据可使编辑结果显著提升,减少伪影和抖动
- 适合视频生成、编辑研究者使用,尤其关注端到端方法
近期视频生成技术的发展推动了视频编辑技术的进步,现有方法可分为基于反演和端到端两类。前者虽无需训练且灵活,但推理耗时长,难以处理细粒度指令,且易产生伪影与抖动;后者依赖成对编辑视频训练,推理快,但因缺乏高质量训练数据而效果不佳。为此,本文提出Señorita-2M,一个包含约200万对高质量视频编辑样本的数据集。该数据集由团队自研并训练的4个专业化视频编辑模型生成,并采用过滤流程剔除低质量样本。我们还评估了主流视频编辑架构,结合预训练生成模型确定最优结构。大量实验表明,该数据集能显著提升视频编辑质量。更多信息见https://senorita-2m-dataset.github.io。
原文摘要 · Abstract (English)
Recent advancements in video generation have spurred the development of video editing techniques, which can be divided into inversion-based and end-to-end methods. However, current video editing methods still suffer from several challenges. Inversion-based methods, though training-free and flexible, are time-consuming during inference, struggle with fine-grained editing instructions, and produce artifacts and jitter. On the other hand, end-to-end methods, which rely on edited video pairs for training, offer faster inference speeds but often produce poor editing results due to a lack of high-quality training video pairs. In this paper, to close the gap in end-to-end methods, we introduce Señorita-2M, a high-quality video editing dataset. Señorita-2M consists of approximately 2 millions of video editing pairs. It is built by crafting four high-quality, specialized video editing models, each crafted and trained by our team to achieve state-of-the-art editing results. We also propose a filtering pipeline to eliminate poorly edited video pairs. Furthermore, we explore common video editing architectures to identify the most effective structure based on current pre-trained generative model. Extensive experiments show that our dataset can help to yield remarkably high-quality video editing results. More details are available at https://senorita-2m-dataset.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。