双向流匹配让少步图像编辑更准更快,不依赖预训练生成器。
BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generation
- 统一建模生成与反演,双向估计图像到噪声和噪声到图像的速度场。
- 少步采样下图像编辑质量显著提升,优于现有方法,支持单步反演。
- 适配主流扩散模型,无需额外模块,可直接集成到现有框架中。
最近的扩散模型和流匹配模型通过迭代去噪实现强大的图像生成与编辑能力。然而,少步采样时前向过程近似较差,导致编辑质量下降。现有少步反演方法常依赖预训练生成器和辅助模块,限制了跨架构的可扩展性与泛化能力。为此,我们提出BiFM(双向流匹配),一种统一框架,在单一模型中联合学习生成与反演。BiFM直接估计图像→噪声和噪声→图像两个方向的平均速度场,受共享瞬时速度场约束,该速度场来自预定义调度或预训练多步扩散模型。此外,BiFM引入连续时间区间监督的新型训练策略,通过双向一致性目标与轻量级时间区间嵌入实现稳定训练。该双向结构还支持单步反演,并能无缝集成至主流扩散与流匹配主干网络。在多种图像编辑与生成任务中,BiFM持续优于现有少步方法,展现出更优性能与更强可编辑性。
原文摘要 · Abstract (English)
Recent diffusion and flow matching models have demonstrated strong capabilities in image generation and editing by progressively removing noise through iterative sampling. While this enables flexible inversion for semantic-preserving edits, few-step sampling regimes suffer from poor forward process approximation, leading to degraded editing quality. Existing few-step inversion methods often rely on pretrained generators and auxiliary modules, limiting scalability and generalization across different architectures. To address these limitations, we propose BiFM (Bidirectional Flow Matching), a unified framework that jointly learns generation and inversion within a single model. BiFM directly estimates average velocity fields in both ``image $\to$ noise" and ``noise $\to$ image" directions, constrained by a shared instantaneous velocity field derived from either predefined schedules or pretrained multi-step diffusion models. Additionally, BiFM introduces a novel training strategy using continuous time-interval supervision, stabilized by a bidirectional consistency objective and a lightweight time-interval embedding. This bidirectional formulation also enables one-step inversion and can integrate seamlessly into popular diffusion and flow matching backbones. Across diverse image editing and generation tasks, BiFM consistently outperforms existing few-step approaches, achieving superior performance and editability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。