arXiv:2508.08134cs.CV2025-08中稿 · ICLR被引 27

无需训练即可精准编辑物体形状,保持背景不变。

Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control

  • 通过比较反演与去噪路径的差异,定位可编辑区域。
  • 在120张图像上实现更高精度的形状替换和画面质量。
  • 适合需要大尺度形变且不希望背景受损的编辑任务。

尽管最近基于流的图像编辑模型具备通用能力,但在涉及大规模形状变换的复杂场景中仍表现不佳,常出现目标形状未正确改变或非目标区域被意外修改的问题。本文提出 Follow-Your-Shape,一种无需训练、无需掩码的框架,可在严格保留非目标内容的前提下,实现精确可控的对象形状编辑。受反演路径与编辑路径间轨迹差异的启发,我们通过对比逐标记速度差计算出轨迹发散图(TDM),精准定位可编辑区域,并引导分阶段的键值注入机制,确保编辑过程稳定且忠实。为支持严谨评估,我们构建了 ReShapeBench,包含120张新图像和精心设计的提示对,专用于形状感知编辑。实验表明,该方法在需大规模形状替换的任务中显著提升可编辑性与视觉保真度。

原文摘要 · Abstract (English)

While recent flow-based image editing models demonstrate general-purpose capabilities across diverse tasks, they often struggle to specialize in challenging scenarios -- particularly those involving large-scale shape transformations. When performing such structural edits, these methods either fail to achieve the intended shape change or inadvertently alter non-target regions, resulting in degraded background quality. We propose Follow-Your-Shape, a training-free and mask-free framework that supports precise and controllable editing of object shapes while strictly preserving non-target content. Motivated by the divergence between inversion and editing trajectories, we compute a Trajectory Divergence Map (TDM) by comparing token-wise velocity differences between the inversion and denoising paths. The TDM enables precise localization of editable regions and guides a Scheduled KV Injection mechanism that ensures stable and faithful editing. To facilitate a rigorous evaluation, we introduce ReShapeBench, a new benchmark comprising 120 new images and enriched prompt pairs specifically curated for shape-aware editing. Experiments demonstrate that our method achieves superior editability and visual fidelity, particularly in tasks requiring large-scale shape replacement.

图像编辑形状控制无训练轨迹分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。