arXiv:2602.03826cs.CVcs.GR2026-02International Conf…被引 4

让图像视频编辑的修改强度可连续调节,效果更平滑自然。

Continuous Control of Editing Models via Adaptive-Origin Guidance

  • 用自适应引导原点动态调整无条件预测,实现渐进式编辑。
  • 在图像和视频编辑任务中,控制过渡更平滑,优于现有滑块方案。
  • 无需额外训练或专用数据集,推理时即可精细调节编辑程度。

基于扩散模型的编辑工具已成为语义图像和视频操作的强大手段。然而,现有模型缺乏对文本引导编辑强度的平滑控制机制。标准的文本条件生成中,无分类器指导(CFG)影响提示遵循度,暗示其可用于编辑强度控制。但本文发现,在这些模型中增大CFG并不能实现输入与编辑结果之间的平滑过渡。我们归因于无条件预测作为引导原点,在低引导尺度下主导生成过程,同时代表对输入内容的任意篡改。为实现连续控制,我们提出自适应原点引导(AdaOr),通过身份条件化的自适应原点替代标准无条件预测,并根据编辑强度对二者进行插值,确保从输入到编辑结果的连续过渡。我们在图像和视频编辑任务上评估了该方法,结果表明其相比当前基于滑块的编辑方式提供了更平滑、更一致的控制效果。该方法将身份指令融入标准训练框架,可在推理阶段实现细粒度控制,无需每项编辑单独处理或依赖特殊数据集。

原文摘要 · Abstract (English)

Diffusion-based editing models have emerged as a powerful tool for semantic image and video manipulation. However, existing models lack a mechanism for smoothly controlling the intensity of text-guided edits. In standard text-conditioned generation, Classifier-Free Guidance (CFG) impacts prompt adherence, suggesting it as a potential control for edit intensity in editing models. However, we show that scaling CFG in these models does not produce a smooth transition between the input and the edited result. We attribute this behavior to the unconditional prediction, which serves as the guidance origin and dominates the generation at low guidance scales, while representing an arbitrary manipulation of the input content. To enable continuous control, we introduce Adaptive-Origin Guidance (AdaOr), a method that adjusts this standard guidance origin with an identity-conditioned adaptive origin, using an identity instruction corresponding to the identity manipulation. By interpolating this identity prediction with the standard unconditional prediction according to the edit strength, we ensure a continuous transition from the input to the edited result. We evaluate our method on image and video editing tasks, demonstrating that it provides smoother and more consistent control compared to current slider-based editing approaches. Our method incorporates an identity instruction into the standard training framework, enabling fine-grained control at inference time without per-edit procedure or reliance on specialized datasets.

图像编辑扩散模型连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。