arXiv:2604.07958cs.CV2026-04

用1.3万张图像对训练视频编辑模型,保留时序一致性且效率高。

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

论文配图:ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
图 1 · 摘自论文原文
  • 基于图像对学习空间变化,冻结3D注意力模块保时序动态。
  • 仅用5轮训练、13K图像对即达主流模型效果。
  • 文本引导动态语义门控实现灵活精准编辑,适合内容创作场景。

当前视频编辑模型多依赖昂贵的成对视频数据,制约其可扩展性。本质上看,多数视频编辑任务可解耦为时空过程:保持预训练模型的时间动态性,仅选择性精确修改空间内容。为此,我们提出ImVideoEdit,一种完全基于图像对学习视频编辑能力的高效框架。通过冻结预训练的3D注意力模块,并将图像视为单帧视频,解耦2D空间学习以保留原始时序动态。核心是预测-更新空间差异注意力模块,逐步提取并注入空间差异。不依赖固定外部掩码,而是引入文本引导的动态语义门控机制,实现自适应、隐式的文本驱动修改。尽管仅在13,000张图像对上训练5个周期,且计算开销极低,ImVideoEdit仍达到与大规模视频数据训练的大型模型相当的编辑保真度与时序一致性。

原文摘要 · Abstract (English)

Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoupled spatiotemporal process, where the temporal dynamics of the pretrained model are preserved while spatial content is selectively and precisely modified. Based on this insight, we propose ImVideoEdit, an efficient framework that learns video editing capabilities entirely from image pairs. By freezing the pre-trained 3D attention modules and treating images as single-frame videos, we decouple the 2D spatial learning process to help preserve the original temporal dynamics. The core of our approach is a Predict-Update Spatial Difference Attention module that progressively extracts and injects spatial differences. Rather than relying on rigid external masks, we incorporate a Text-Guided Dynamic Semantic Gating mechanism for adaptive and implicit text-driven modifications. Despite training on only 13K image pairs for 5 epochs with exceptionally low computational overhead, ImVideoEdit achieves editing fidelity and temporal consistency comparable to larger models trained on extensive video datasets.

视频编辑图像对训练时空解耦文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。