用视频预训练的时序先验,让图像编辑仅需1%数据量就能达到顶尖效果。
Video4Edit: Viewing Image Editing as a Degenerate Temporal Process
- 把图像编辑看作退化的时序过程,复用视频模型的演化先验。
- 仅用主流方法1%的标注数据,性能接近领先开源模型。
- 适合追求数据效率和低资源部署的图像编辑研究者。
我们观察到,多模态基础模型的进步已使指令驱动的图像生成与编辑进入真正的跨模态协作阶段。然而,当前最先进的编辑流程仍成本高昂:除了训练大型扩散/流模型外,还需收集大量高质量三元组({指令, 原图, 编辑后图})以覆盖多样用户意图。此外,视觉替换的保真度取决于指令对目标语义的精准描述。本文从时序建模视角重新审视这一挑战:若视频可视为完整的时序过程,则图像编辑可被视为退化的时序过程。该视角使我们能够迁移视频预训练中的单帧演化先验,实现高度数据高效的微调。实验表明,我们的方法在仅使用主流编辑模型约1%监督数据的情况下,性能与领先开源基线相当。
原文摘要 · Abstract (English)
We observe that recent advances in multimodal foundation models have propelled instruction-driven image generation and editing into a genuinely cross-modal, cooperative regime. Nevertheless, state-of-the-art editing pipelines remain costly: beyond training large diffusion/flow models, they require curating massive high-quality triplets of \{instruction, source image, edited image\} to cover diverse user intents. Moreover, the fidelity of visual replacements hinges on how precisely the instruction references the target semantics. We revisit this challenge through the lens of temporal modeling: if video can be regarded as a full temporal process, then image editing can be seen as a degenerate temporal process. This perspective allows us to transfer single-frame evolution priors from video pre-training, enabling a highly data-efficient fine-tuning regime. Empirically, our approach matches the performance of leading open-source baselines while using only about one percent of the supervision demanded by mainstream editing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。