arXiv:2509.21905cs.CV2025-09被引 1

融合文字与拖拽的统一图像编辑框架,实现精准空间控制与纹理细节同步调整。

TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation

  • 通过3D特征映射实现确定性拖拽,提升潜在空间布局控制力。
  • 动态平衡文本与拖拽影响,支持联合、纯文本或纯拖拽多种模式。
  • 在多类编辑任务中表现优于专用模型,适合需要灵活控制的生成应用。

本文探索在文字与拖拽交互共同控制下的图像编辑。尽管近期文字驱动和拖拽驱动的编辑方法已取得显著进展,但各自存在互补性局限:文字驱动方法擅长纹理修改,却缺乏精确的空间控制;而拖拽方法主要改变形状与结构,缺少细粒度纹理引导。为此,我们提出一种统一的基于扩散模型的联合拖拽-文字图像编辑框架,融合两种范式优势。本框架引入两项关键创新:(1) 点云确定性拖拽,通过三维特征映射增强潜在空间布局控制;(2) 拖拽-文字引导去噪,在去噪过程中动态平衡拖拽与文本条件的影响。值得注意的是,该模型支持灵活的编辑模式——仅使用文字、仅使用拖拽,或两者结合——且在每种模式下均保持优异性能。大量定量与定性实验表明,该方法不仅实现了高保真联合编辑,还达到或超越了专用文字或拖拽方法的性能,为可控图像操作提供了一个通用且可推广的解决方案。代码将公开以复现所有结果。

原文摘要 · Abstract (English)

This paper explores image editing under the joint control of text and drag interactions. While recent advances in text-driven and drag-driven editing have achieved remarkable progress, they suffer from complementary limitations: text-driven methods excel in texture manipulation but lack precise spatial control, whereas drag-driven approaches primarily modify shape and structure without fine-grained texture guidance. To address these limitations, we propose a unified diffusion-based framework for joint drag-text image editing, integrating the strengths of both paradigms. Our framework introduces two key innovations: (1) Point-Cloud Deterministic Drag, which enhances latent-space layout control through 3D feature mapping, and (2) Drag-Text Guided Denoising, dynamically balancing the influence of drag and text conditions during denoising. Notably, our model supports flexible editing modes - operating with text-only, drag-only, or combined conditions - while maintaining strong performance in each setting. Extensive quantitative and qualitative experiments demonstrate that our method not only achieves high-fidelity joint editing but also matches or surpasses the performance of specialized text-only or drag-only approaches, establishing a versatile and generalizable solution for controllable image manipulation. Code will be made publicly available to reproduce all results presented in this work.

图像编辑扩散模型交互控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。