0.23秒完成文本引导图像编辑,速度比之前快50倍
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
- 一步逆向重建图像,避免多步采样耗时
- 0.23秒实现编辑,比多步方法快至少50倍
- 适合实时应用与设备端部署
近期的文本引导图像编辑技术利用多步扩散模型的丰富先验,通过简单文本输入实现图像修改。然而,由于涉及昂贵的多步反演和采样过程,这些方法难以满足真实场景和设备端应用的速度需求。为此,我们提出SwiftEdit,一种高效且简洁的编辑工具,可在0.23秒内实现即时文本引导图像编辑。其核心贡献包括:一种一步逆向框架,通过一次逆向实现图像重建;以及基于掩码的编辑技术,结合提出的注意力重缩放机制,实现局部图像修改。大量实验表明,SwiftEdit在保持良好编辑效果的同时,实现远超以往多步方法的效率,最快可达50倍以上加速。项目主页:https://swift-edit.github.io/
原文摘要 · Abstract (English)
Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applications due to the costly multi-step inversion and sampling process involved. In response to this, we introduce SwiftEdit, a simple yet highly efficient editing tool that achieve instant text-guided image editing (in 0.23s). The advancement of SwiftEdit lies in its two novel contributions: a one-step inversion framework that enables one-step image reconstruction via inversion and a mask-guided editing technique with our proposed attention rescaling mechanism to perform localized image editing. Extensive experiments are provided to demonstrate the effectiveness and efficiency of SwiftEdit. In particular, SwiftEdit enables instant text-guided image editing, which is extremely faster than previous multi-step methods (at least 50 times faster) while maintain a competitive performance in editing results. Our project page is at: https://swift-edit.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。