arXiv:2608.26101cs.CV2026-08

构建600万样本参考视频数据集,提升视频编辑的准确性与可控性。

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

论文配图:RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
图 1 · 摘自论文原文
  • 用真实无瑕疵视频作目标,通过多专家过滤生成高质量输入条件。
  • 包含约600万视觉参考,支持基于图像与视频的精细对应学习。
  • 适配需要高保真、身份一致编辑的研究者与开发者使用。

近期视频编辑进展主要依赖大规模指令驱动数据集,但现有数据集仍存在两大缺陷:一是目标视频常由自动编辑模型生成,易引入可见伪影和不可靠监督信号;二是多数公开数据集仅依赖文本指令,缺乏对精确、身份保持及可控编辑至关重要的视觉参考。为此,我们提出RefVideo-6M,一个大规模参考引导式编辑数据集,包含500万视频编辑样本和100万图像编辑样本。为确保可靠监督,数据集采用处理无伪影真实视频为目标、通过多编辑专家生成质量过滤后的输入条件。此外,提供约600万视觉参考,涵盖多样参考类型与编辑场景,使模型能学习超越纯文本的细粒度视觉对应关系。基于RefVideo-6M,我们进一步训练了参考引导式视频编辑模型Ref-MoT,以评估数据集的有效性与可扩展性。大量实验表明,相比现有数据集,RefVideo-6M提供更可靠的监督信号,并支持训练出在视觉质量、可控性和参考一致性方面均有提升的强大编辑模型。开源数据集已发布于https://huggingface.co/datasets/RefVideo6M/RefVideo6M。

原文摘要 · Abstract (English)

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by automatic editing models, which may introduce visible artifacts and unreliable supervision signals. Second, most public datasets rely primarily on textual instructions, while lacking visual references that are crucial for precise, identity-preserving, and controllable editing. To address these limitations, we introduce RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples. To ensure reliable supervision, our dataset uses a construction pipeline that treats artifact-free real videos as editing targets and generates quality-filtered input conditions with multiple editing experts. In addition, it provides approximately 6 million visual references, covering diverse reference types and editing scenarios, thereby enabling models to learn fine-grained visual correspondence beyond text-only instructions. Based on RefVideo-6M, we further train a reference-guided video editing model, Ref-MoT, to evaluate the effectiveness and scalability of the proposed dataset. Extensive experiments demonstrate that RefVideo-6M provides substantially more reliable supervision than existing datasets and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency. The open-source dataset is available at https://huggingface.co/datasets/RefVideo6M/RefVideo6M.

视频编辑参考学习数据集真实视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。