arXiv:2411.16819cs.CVcs.AI2024-11CVPR被引 30

用视频生成模型做图像编辑,让修改更自然精准。

Pathways on the Image Manifold: Image Editing via Video Generation

  • 把图像编辑看作时间序列,用预训练视频模型生成连续过渡
  • 在文本编辑任务上达到当前最优,显著提升准确率与原图保真度
  • 适合需要精细控制和高保真的图像编辑场景

近期基于图像扩散模型的图像编辑技术取得显著进展,但仍面临难以准确遵循复杂编辑指令、常牺牲原图关键元素保真度的问题。与此同时,视频生成模型已发展为能持续模拟真实世界的一致性系统。本文提出将二者结合,利用图像到视频模型进行图像编辑。将图像编辑重构为时间过程,通过预训练视频模型从原始图像平滑过渡到目标编辑结果。该方法沿图像流形连续遍历,确保编辑一致性并保留原图核心特征。实验表明,该方法在文本驱动的图像编辑任务中达到当前最优性能,显著提升编辑准确率与图像保真度。

原文摘要 · Abstract (English)

Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently compromise fidelity by altering key elements of the original image. Simultaneously, video generation has made remarkable strides, with models that effectively function as consistent and continuous world simulators. In this paper, we propose merging these two fields by utilizing image-to-video models for image editing. We reformulate image editing as a temporal process, using pretrained video models to create smooth transitions from the original image to the desired edit. This approach traverses the image manifold continuously, ensuring consistent edits while preserving the original image's key aspects. Our approach achieves state-of-the-art results on text-based image editing, demonstrating significant improvements in both edit accuracy and image preservation. Visit our project page at https://rotsteinnoam.github.io/Frame2Frame.

图像编辑视频生成扩散模型连续编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。