arXiv:2607.10800cs.CV2026-07

用数学变换实现无需训练的精准图像编辑

h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform

论文配图:h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform
图 1 · 摘自论文原文
  • 基于杜布变换构建条件生成框架,统一源图一致与目标对齐
  • 实现闭式重建引导与语义编辑信号,支持可控权衡
  • 无需训练、鲁棒性强,适合各类图像编辑场景

使用预训练文本到图像流模型进行图像编辑通常需要在目标对齐与源图一致性之间精细平衡。现有方法或依赖反演流程,或采用启发式源到目标轨迹构造,常受限于特定架构设计或对超参数敏感。本文提出 h-Flow,一种无训练且理论严谨的流式图像编辑框架。受杜布 $h$-变换启发,我们将图像编辑重构为多个终端事件下的条件生成问题,对应源图一致性和目标对齐性。我们首先将经典的 $h$-变换从SDE模型扩展至确定性RF框架,通过构建具有相同边缘分布的等效SDE。在此框架下,设计了针对源一致性与目标对齐性的专用 $h$-函数,获得闭式重建引导和基于速度的语义编辑信号。进一步引入速度正交分解,解耦重建与编辑方向,实现两者间的可控权衡。大量实验表明,h-Flow在多种场景下均能实现高效、鲁棒且灵活的编辑效果。代码即将发布。

原文摘要 · Abstract (English)

Editing images with pre-trained text-to-image flow models typically requires carefully balancing target alignment with the desired prompt and source consistency with the original image. Existing approaches either rely on inversion-based pipelines or heuristic source-to-target trajectory constructions, which often depend on architecture-specific designs or are sensitive to hyperparameters. In this paper, we propose h-Flow, a training-free and theoretically grounded flow-based editing framework. Inspired by Doob's $h$-Transform, we reformulate image editing as conditional generation under multiple terminal events corresponding to source consistency and target alignment. We first extend the classical $h$-Transform from SDE-based models to the deterministic RF framework by constructing an equivalent SDE with identical marginals. Within this formulation, we design dedicated $h$-functions for source consistency and target alignment, yielding closed-form reconstruction guidance and velocity-based semantic editing signals. We further introduce a velocity orthogonal decomposition to decouple reconstruction and editing directions, enabling a controllable trade-off between the two objectives. Extensive experiments demonstrate that h-Flow achieves effective, robust, and flexible editing across diverse scenarios. The code will be released soon.

图像编辑流模型无训练概率变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。