arXiv:2503.17641cs.CV2025-03被引 7

构建全流程视频指令编辑框架,解决数据稀缺与泛化难题

InstructVEdit: A Holistic Approach for Instructional Video Editing

  • 设计数据清洗流程,构建可靠训练数据集
  • 改进模型结构,提升编辑质量与时间一致性
  • 采用迭代优化策略,增强真实场景适应性

根据指令进行视频编辑是一项极具挑战性的任务,主要受限于高质量编辑视频对数据的稀缺。这种数据匮乏不仅制约了训练数据的获取,也阻碍了模型架构与训练策略的系统性探索。尽管已有研究在特定环节有所改进(如利用图像编辑技术生成视频数据集或分解视频编辑训练),但针对上述问题的完整框架仍鲜有涉及。本文提出InstructVEdit,一种端到端的指令式视频编辑方法:(1) 建立可靠的训练数据集构建流程;(2) 引入两种模型架构优化,以提升编辑质量并保持时间一致性;(3) 提出基于真实数据的迭代精炼策略,增强泛化能力并减少训练与测试间的差异。大量实验表明,InstructVEdit在指令驱动视频编辑任务中达到当前最优性能,展现出对多样化真实场景的强适应能力。

原文摘要 · Abstract (English)

Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the systematic exploration of model architectures and training strategies. While prior work has improved specific aspects of video editing (e.g., synthesizing a video dataset using image editing techniques or decomposed video editing training), a holistic framework addressing the above challenges remains underexplored. In this study, we introduce InstructVEdit, a full-cycle instructional video editing approach that: (1) establishes a reliable dataset curation workflow to initialize training, (2) incorporates two model architectural improvements to enhance edit quality while preserving temporal consistency, and (3) proposes an iterative refinement strategy leveraging real-world data to enhance generalization and minimize train-test discrepancies. Extensive experiments show that InstructVEdit achieves state-of-the-art performance in instruction-based video editing, demonstrating robust adaptability to diverse real-world scenarios. Project page: https://o937-blip.github.io/InstructVEdit.

视频编辑指令生成模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。