arXiv:2512.03981cs.CV2025-12被引 1

无需遮罩和提示词,仅用拖动即可高保真编辑图像。

DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment

  • 自动推断可编辑区域,基于点位移生成软遮罩。
  • 利用扩散模型中间激活值对齐特征,保持结构一致性。
  • 在无任何人工输入下仍实现高质量图像编辑,适合交互设计场景。

基于生成模型的拖拽编辑提供了对图像结构的直观控制。然而,现有方法严重依赖手动提供的遮罩和文本提示以保证语义保真度与运动精度。去除这些约束会导致根本性权衡:缺乏遮罩时出现视觉伪影,缺少提示则空间控制不佳。为此,我们提出DirectDrag,一种全新的无遮罩、无提示编辑框架。DirectDrag通过最少用户输入实现精确高效的操控,同时保持高图像保真度和准确的点对齐。其核心创新包括:首先,设计了自动软遮罩生成模块,从点位移智能推断可编辑区域,沿移动路径自动定位形变,同时通过生成模型固有的能力保留上下文完整性;其次,开发读出引导特征对齐机制,利用扩散模型中间激活值维持点编辑过程中的结构一致性,显著提升视觉质量。尽管不依赖手动遮罩或提示,DirectDrag在图像质量上优于现有方法,且保持了竞争性的拖拽准确性。在DragBench及真实场景中的大量实验验证了DirectDrag在高质量交互式图像编辑中的有效性与实用性。

原文摘要 · Abstract (English)

Drag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask- and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model's inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation. Project Page: https://frakw.github.io/DirectDrag/. Code is available at: https://github.com/frakw/DirectDrag.

图像编辑扩散模型交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。