arXiv:2609.06288cs.CV2026-09

无需训练,精准控制物体编辑时的背景稳定性和边界清晰度

Object-Aware Background-Controlled Editing via Weighted Velocity Guidance

论文配图:Object-Aware Background-Controlled Editing via Weighted Velocity Guidance
图 1 · 摘自论文原文
  • 分离语义残差作用位置与注入方式,实现对象级控制
  • 通过约束注入抑制背景漂移,时间自适应加权稳定边界过渡
  • 适用于图像视频编辑,保持局部可编辑性且无模型微调

无需训练的图像编辑在推理阶段通过修改提示条件下的去噪速度来操控生成模型。现有基于速度的编辑方法通常将提示引起的残差全局作用于潜在空间,并依赖模型隐式定位语义变化。在以物体为中心的编辑中,这些残差在目标物体外很少为零,导致多步积分时小的非目标成分累积,引发背景漂移和边界不稳定。我们提出训练自由的物体感知速度控制(OAVC),在速度积分过程中引入物体级控制。OAVC 将语义残差的作用位置与注入方式解耦:在源提示下构建背景锚定的参考界面,再在目标提示下进行物体局部的安全语义注入。受限注入算子抑制导致漂移的速度分量,而时间自适应空间加权则稳定物体边缘处的过渡。OAVC 不需要训练或修改预训练模型参数。在基于图像和视频修正流(rectified-flow)骨干网络的对象中心图像与视频基准测试中,实验显示其在保留有效局部可编辑性的前提下,显著提升了背景保持、结构保真度、边界稳定性和时间一致性。

原文摘要 · Abstract (English)

Training-free image editing steers diffusion or flow-matching generative models at inference time by modifying prompt-conditioned denoising velocities. Existing velocity-based editors often apply prompt-induced residuals globally over the latent space and rely on the model to localize semantic changes implicitly. For object-centric edits, these residuals are rarely zero outside the target object, so small non-target components can accumulate during multi-step integration, causing background drift and unstable object boundaries. We propose Object-Aware Velocity Control (OAVC), a training-free framework that introduces object-level control into the velocity-integration process. OAVC decouples where semantic residuals are allowed to act from how they are injected into the dynamics. It constructs a background-anchored reference interface under the source prompt and then performs object-localized safe semantic injection under the target prompt. A constrained injection operator suppresses drift-inducing velocity components, while time-adaptive spatial weighting stabilizes the transition near object boundaries. OAVC requires no training or modification of pretrained model parameters. Experiments on object-centric image and video benchmarks with image and video rectified-flow backbones show improved background preservation, structural fidelity, boundary stability, and temporal consistency while retaining effective localized editability.

图像编辑扩散模型速度控制无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。