用图像编辑数据训练视频编辑模型,提升效果与多样性
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

- 融合大规模图像编辑数据与视频数据,联合训练视频编辑模型
- 通过帧级噪声策略缓解图像与视频的领域差异,生成合理时序变化
- 构建60+挑战性任务评测基准,适合研究视频生成与编辑的学者
视频编辑旨在根据用户意图修改输入视频。近年来,端到端训练方法受到广泛关注,通过视频生成或编辑模型构建配对视频编辑数据。然而,相较于图像编辑,视频数据的高标注成本严重制约了视频编辑数据集的规模、质量和任务多样性。为此,我们提出LIVE,一种联合训练框架,利用大规模高质量图像编辑数据与视频数据共同增强编辑能力。为缓解静态图像与动态视频之间的领域差异,我们引入帧级令牌噪声策略,将特定帧的潜在表示视为推理令牌,借助大预训练视频生成模型生成合理的时序变换。此外,通过清洗公开数据集并构建自动化数据流水线,采用两阶段训练策略逐步提升视频编辑能力。同时,我们构建了一个涵盖60余项常见于图像编辑但稀缺于现有视频数据集的综合性评估基准。大量对比与消融实验表明,该方法达到当前最佳性能。源代码将公开。
原文摘要 · Abstract (English)
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing models. However, compared to image editing, the high annotation costs of video data severely constrain the scale, quality, and task diversity of video editing datasets when relying on video generative models or manual annotation. To bridge this gap, we propose LIVE, a joint training framework that leverages large-scale, high-quality image editing data alongside video datasets to bolster editing capabilities. To mitigate the domain discrepancy between static images and dynamic videos, we introduce a frame-wise token noise strategy, which treats the latents of specific frames as reasoning tokens, leveraging large pretrained video generative models to create plausible temporal transformations. Moreover, through cleaning public datasets and constructing an automated data pipeline, we adopt a two-stage training strategy to anneal video editing capabilities. Furthermore, we curate a comprehensive evaluation benchmark encompassing over 60 challenging tasks that are prevalent in image editing but scarce in existing video datasets. Extensive comparative and ablation experiments demonstrate that our method achieves state-of-the-art performance. The source code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。