无需遮罩或提示词,拖拽即可实时生成逼真图像。
InstantDrag: Improving Interactivity in Drag-based Image Editing
- 用光流生成与扩散模型分步处理拖拽动作
- 在人脸和通用场景上实现秒级真实图像编辑
- 适合需要快速交互的视觉设计与实时应用
基于拖拽的图像编辑因其交互性和精准性而受到关注。然而,尽管文生图模型能在一秒内生成样本,拖拽编辑仍因难以准确反映用户操作且保持图像内容而滞后。现有方法依赖计算量大的逐图优化或复杂的引导机制,需额外输入如可移动区域掩码或文本提示,损害了交互性。本文提出InstantDrag,一种无优化流程的方案,仅需图像和拖拽指令作为输入。InstantDrag包含两个精心设计的网络:拖拽条件光流生成器(FlowGen)与光流条件扩散模型(FlowDiffusion)。通过分解任务为运动生成与运动条件图像生成,该方法在真实视频数据集上学习拖拽动态。实验表明,InstantDrag可在无需掩码或文本提示的情况下实现快速、逼真的图像编辑,适用于人脸视频数据集与通用场景。结果凸显了该方法在拖拽编辑中的高效性,为交互式实时应用提供了有前景的解决方案。
原文摘要 · Abstract (English)
Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。