构建百万级视频编辑数据集,支持复杂指令控制。
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

- 设计分解式合成流程与渐进过滤系统生成高质量数据。
- 提出Goku-Edit模型,实现指令遵循提升8%。
- 适合研究多任务视频编辑与大模型应用的开发者。
现有基于指令的视频编辑数据集多聚焦于单一任务的外观修改,难以满足真实场景下的复杂创作需求。为此,我们提出Goku,首个涵盖200万条高质量、指令对齐的视频编辑样本的数据集,首次将任务范围从基础外观编辑扩展至多任务与结构化操作(如精确控制主体运动)。为应对复杂任务的数据合成挑战,我们设计了一种高效的合成流程,将复杂编辑分解为可控制的子问题,并引入全链路渐进式过滤系统以保障数据可靠性。此外,我们在Goku上探索最优网络结构,提出Goku-Edit模型。为深入理解复杂编辑指令,Goku-Edit采用多模态大模型作为文本编码器,并采用解耦双分支架构:专用掩码分支负责结构控制,主分支专注外观渲染。同时,我们构建了包含1,000个经人工验证的测试用例和7项新型编辑专用评估指标的Goku-Bench基准。在Goku-Bench上评估显示,Goku-Edit相较其他开源模型在指令遵循能力上最高提升8%。
原文摘要 · Abstract (English)
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge this gap, we present Goku, a large-scale dataset featuring 2 million high-quality, instruction-aligned video editing pairs, which is the first to extend task boundaries from basic appearance editing to multi-task and structural manipulations(e.g., precise control of subject movement). To tackle the data synthesis challenges inherent in these complex tasks, we design an efficient data synthesis pipeline that decomposes complex edits into controllable sub-problems and introduce a progressive filtering system for data reliability throughout the whole process. Furthermore, we explore the optimal network structures on Goku, and propose Goku-Edit. To deeply comprehend complex editing instructions, Goku-Edit leverages an MLLM as its text encoder and adopts a decoupled dual-branch design: a dedicated mask branch handles structural control, freeing the main branch for appearance rendering. A comprehensive video editing benchmark, Goku-Bench, is also proposed with 1,000 human-verified test cases and 7 novel editing-specific metrics. Evaluated on Goku-Bench, Goku-Edit obtains up to +8% improvement on other open-source models in terms of instruction following.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。