arXiv:2605.13349cs.CV2026-05

通过约束潜在空间分布,实现更自然的文本引导图像点编辑

Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints

论文配图:Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints
图 1 · 摘自论文原文
  • 用CLIP评估中间步骤,确保语义一致性
  • 引入先验保持损失,控制潜在变量不偏离原始分布
  • 方向加权追踪提升精度与效率,适合精细编辑

基于扩散模型的点编辑方法因能通过噪声潜空间的局部扰动操控图像语义和细节而受到关注。但传统方法依赖手柄点与目标点定义运动轨迹,易引发歧义或非必要修改;当两点距离较远时,累积扰动会导致潜在变量偏离反演得分轨迹,产生不自然伪影。为此,我们提出基于CLIP的模型评估并指导中间编辑步骤,确保生成结果语义一致。同时设计先验保持损失,约束优化后的潜在代码保持在扩散先验的采样空间内,使生成沿熟悉得分轨迹进行。针对细粒度任务,提出方向加权点追踪机制,在相似特征区域内引导编辑方向,提升追踪精度与生成质量,同时减少编辑时间。

原文摘要 · Abstract (English)

Diffusion-based point editing methods have gained significant traction in image editing tasks due to their ability to manipulate image semantics and fine details by applying localized perturbations on the manifold of noise latent. However, these approaches face several limitations. Traditional point-based editing relies on pairs of handle and target points to define motion trajectories, which can introduce ambiguity or unnecessary alterations. Furthermore, when the distance between the handle and target points is large, the accumulated perturbations often cause the noise latent deviation from inversion score trajectory, resulting in unnatural artifacts. To address these issues in global editing tasks, we introduce a CLIP-based model to evaluate and guide intermediate editing steps, ensuring that the generated results remain both semantically aligned. Additionally, we propose a prior-preservation loss that constrains the optimized latent code to stay within the sampling space of the diffusion prior, improving consistency with the original data distribution, to ensure the model generates images along a familiar score trajectory. For fine-grained tasks, we present a directionally-weighted point tracking mechanism that steers the editing process toward the target direction within similar feature regions. This improves both the tracking accuracy and generation quality, while also reducing the editing time.

图像编辑扩散模型点追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。