解决流模型编辑中的轨迹漂移问题,实现精准语义转换。
On Exact Editing of Flow-Based Diffusion Models
- 引入双视角速度转换机制,分离结构保持与语义引导分支。
- 通过后验一致更新补偿累积速度误差,提升轨迹稳定性。
- 适用于需要高保真图像编辑的科研与工业场景。
基于流的扩散模型编辑方法可直接实现源图像分布到目标图像分布的转换,无需显式反演。然而,这些方法的潜在轨迹常因速度误差累积导致语义不一致和结构失真。本文提出条件速度校正(CVC),将流编辑重构为由已知源先验驱动的分布变换问题。CVC通过双视角速度转换机制,将潜在演化分解为保持结构的分支与引导语义的分支。条件速度场相对于真实分布轨迹存在绝对速度误差,易引发潜空间不稳定与轨迹漂移。为此,我们采用基于经验贝叶斯推断与Tweedie校正的后验一致更新,实现数学上严谨的误差补偿。实验表明,该方法在多种任务中均显著提升保真度、语义对齐性与编辑可靠性,实现稳定且可解释的潜在动态。
原文摘要 · Abstract (English)
Recent methods in flow-based diffusion editing have enabled direct transformations between source and target image distribution without explicit inversion. However, the latent trajectories in these methods often exhibit accumulated velocity errors, leading to semantic inconsistency and loss of structural fidelity. We propose Conditioned Velocity Correction (CVC), a principled framework that reformulates flow-based editing as a distribution transformation problem driven by a known source prior. CVC rethinks the role of velocity in inter-distribution transformation by introducing a dual-perspective velocity conversion mechanism. This mechanism explicitly decomposes the latent evolution into two components: a structure-preserving branch that remains consistent with the source trajectory, and a semantically-guided branch that drives a controlled deviation toward the target distribution. The conditional velocity field exhibits an absolute velocity error relative to the true underlying distribution trajectory, which inherently introduces potential instability and trajectory drift in the latent space. To address this quantifiable deviation and maintain fidelity to the true flow, we apply a posterior-consistent update to the resulting conditional velocity field. This update is derived from Empirical Bayes Inference and Tweedie correction, which ensures a mathematically grounded error compensation over time. Our method yields stable and interpretable latent dynamics, achieving faithful reconstruction alongside smooth local semantic conversion. Comprehensive experiments demonstrate that CVC consistently achieves superior fidelity, better semantic alignment, and more reliable editing behavior across diverse tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。