构建首个大规模高质量视频编辑指令数据集,推动AI视频修改能力发展
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- 设计全流程数据管道,生成高精度视频编辑指令对
- 涵盖7类空间与非空间编辑,共300万条数据,指令平均长度超15词
- 配套基准测试集与50亿参数模型,适合视频生成与编辑研究者使用
指令式图像编辑数据集的质量与多样性持续提升,但指令式视频编辑的大规模高质量数据集仍十分稀缺。为填补这一空白,我们推出了OpenVE-3M——一个开源、大规模、高质量的指令式视频编辑数据集。该数据集包含两类主要编辑类型:空间对齐编辑(全局风格、背景替换、局部修改、局部移除、局部添加、字幕编辑)与非空间对齐编辑(多镜头拍摄编辑、创意编辑)。所有编辑内容均通过精心设计的数据流水线生成,并经过严格质量筛选。OpenVE-3M在规模、编辑类型多样性、指令长度和整体质量上均超越现有开源数据集。此外,为解决领域内缺乏统一评估基准的问题,我们构建了OpenVE-Bench,包含431组视频编辑对,覆盖多样化编辑任务,采用三项与人类判断高度一致的指标进行评测。我们还推出了基于该数据集训练的50亿参数模型OpenVE-Edit,其在OpenVE-Bench上表现卓越,超越所有先前开源模型,包括140亿参数基线。项目主页见 https://lewandofskee.github.io/projects/OpenVE。
原文摘要 · Abstract (English)
The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. It comprises two primary categories: spatially-aligned edits (Global Style, Background Change, Local Change, Local Remove, Local Add, and Subtitles Edit) and non-spatially-aligned edits (Camera Multi-Shot Edit and Creative Edit). All edit types are generated via a meticulously designed data pipeline with rigorous quality filtering. OpenVE-3M surpasses existing open-source datasets in terms of scale, diversity of edit types, instruction length, and overall quality. Furthermore, to address the lack of a unified benchmark in the field, we construct OpenVE-Bench, containing 431 video-edit pairs that cover a diverse range of editing tasks with three key metrics highly aligned with human judgment. We present OpenVE-Edit, a 5B model trained on our dataset that demonstrates remarkable efficiency and effectiveness by setting a new state-of-the-art on OpenVE-Bench, outperforming all prior open-source models including a 14B baseline. Project page is at https://lewandofskee.github.io/projects/OpenVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。