arXiv:2510.14648cs.CVcs.AI2025-10被引 33

用无配对视频训练视频编辑模型,低成本实现精准指令响应

In-Context Learning with Unpaired Clips for Instruction-based Video Editing

论文配图:In-Context Learning with Unpaired Clips for Instruction-based Video Editing
图 1 · 摘自论文原文
  • 通过无配对视频片段实现上下文学习,构建视频编辑基础能力
  • 在100万真实视频上预训练,微调仅需不足15万标注对
  • 相比现有方法提升12%指令遵循度和15%视觉质量,适合高效视频编辑

尽管基于指令的图像编辑发展迅速,但其向视频领域的拓展仍受限于大规模配对视频编辑数据集的高昂成本与复杂性。为此,我们提出一种低成本预训练策略,利用无配对视频片段进行上下文学习,以实现基于指令的视频编辑。实验表明,该策略使基础视频生成模型具备添加、替换或删除等通用编辑能力,可依据输入指令准确执行。预训练后,仅需少量高质量配对数据即可高效微调。基于HunyuanVideoT2V框架,模型先在约100万条真实视频片段上预训练以学习基本编辑概念,再在少于15万条精选编辑对上微调,扩展更多任务并提升质量。对比实验显示,本方法在指令对齐与视觉保真度方面均优于现有方案,指令遵循率提升12%,编辑质量提高15%。

原文摘要 · Abstract (English)

Despite the rapid progress of instruction-based image editing, its extension to video remains underexplored, primarily due to the prohibitive cost and complexity of constructing large-scale paired video editing datasets. To address this challenge, we introduce a low-cost pretraining strategy for instruction-based video editing that leverages in-context learning from unpaired video clips. We show that pretraining a foundation video generation model with this strategy endows it with general editing capabilities, such as adding, replacing, or deleting operations, according to input editing instructions. The pretrained model can then be efficiently refined with a small amount of high-quality paired editing data. Built upon HunyuanVideoT2V, our framework first pretrains on approximately 1M real video clips to learn basic editing concepts, and subsequently fine-tunes on fewer than 150k curated editing pairs to extend more editing tasks and improve the editing quality. Comparative experiments show that our method surpasses existing instruction-based video editing approaches in both instruction alignment and visual fidelity, achieving a 12\% improvement in editing instruction following and a 15\% improvement in editing quality.

视频编辑指令学习无配对数据预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。