用少量数据让视频模型实现精准指令编辑,还能顺便处理图片
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

- 用互信息注意力构建对齐视频对,支持中间帧开始编辑
- 仅需10万级数据即达开源模型顶尖水平
- 训练时混入图像数据,模型自然具备图像编辑能力
基于指令的视频编辑可通过文本控制内容,但将生成模型转化为编辑器通常需要大量数据,而高质量视频编辑数据仍稀缺。本文表明,无需大规模视频编辑数据,视频生成主干模型也能成为强编辑器。我们提出InsEdit,基于HunyuanVideo-1.5构建的指令式编辑模型,结合视觉编辑架构与基于互上下文注意力(MCA)的视频数据流水线,生成对齐视频对,使编辑可从片段中间开始而非仅限首帧。仅使用约10万条视频编辑数据,InsEdit在自建视频指令编辑基准上达到开源方法最优表现。此外,因训练中包含图像编辑数据,模型无需修改即可支持图像编辑。
原文摘要 · Abstract (English)
Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this paper, we show that a video generation backbone can become a strong video editor without large scale video editing data. We present InsEdit, an instruction-based editing model built on HunyuanVideo-1.5. InsEdit combines a visual editing architecture with a video data pipeline based on Mutual Context Attention (MCA), which creates aligned video pairs where edits can begin in the middle of a clip rather than only from the first frame. With only O(100)K video editing data, InsEdit achieves state-of-the-art results among open-source methods on our video instruction editing benchmarks. In addition, because our training recipe also includes image editing data, the final model supports image editing without any modification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。