arXiv:2411.15260cs.CVcs.AI2024-11被引 23

首个大规模视频局部编辑数据集,支持交互式精准修改。

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

  • 基于关键帧引导的迭代编辑机制,实现高效交互。
  • 970万样本数据集,覆盖多种视频编辑任务。
  • 适合视频编辑、AI创意工具开发者使用。

基于扩散的图像编辑模型近年取得显著进展,但高质量视频编辑仍面临挑战。主要瓶颈在于缺乏基于真实数据的大规模开源视频编辑数据集,且视频数据需更多标记符表示,大幅增加训练成本。此外,现有模型交互性差,用户难以一次准确表达编辑意图。为此,本文提出首个大规模混合图像-视频局部编辑数据集 VIVID-10M 和基线模型 VIVID。VIVID-10M 包含 970 万样本,涵盖多样化的视频编辑任务,旨在降低数据构建与模型训练成本。VIVID 模型在该数据集上训练,支持实体添加、修改与删除。核心是关键帧引导的交互式编辑机制,用户可逐帧迭代修改并传播至其他帧,有效降低达成目标的延迟。实验表明,该方法在自动指标与用户评估中均达到当前最优性能。VIVID-10M 数据集已开源:https://kwaivgi.github.io/VIVID/。

原文摘要 · Abstract (English)

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle is the absence of open-source, large-scale video editing datasets based on real-world data, as constructing such datasets is both time-consuming and costly. Moreover, video data requires a significantly larger number of tokens for representation, which substantially increases the training costs for video editing models. Lastly, current video editing models offer limited interactivity, often making it difficult for users to express their editing requirements effectively in a single attempt. To address these challenges, this paper introduces a dataset VIVID-10M and a baseline model VIVID. VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, which comprises 9.7M samples that encompass a wide range of video editing tasks. VIVID is a Versatile and Interactive VIdeo local eDiting model trained on VIVID-10M, which supports entity addition, modification, and deletion. At its core, a keyframe-guided interactive video editing mechanism is proposed, enabling users to iteratively edit keyframes and propagate it to other frames, thereby reducing latency in achieving desired outcomes. Extensive experimental evaluations show that our approach achieves state-of-the-art performance in video local editing, surpassing baseline methods in both automated metrics and user studies. The VIVID-10M dataset are open-sourced at https://kwaivgi.github.io/VIVID/.

视频编辑数据集交互式扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。