arXiv:2508.02363cs.CV2025-08被引 5

用最优传输统一修复与直接编辑,提升图像修改精度与结构保持能力。

Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods

论文配图:Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
图 1 · 摘自论文原文
  • 基于最优传输理论,在逆向扩散中引入零样本引导,改善重建质量。
  • 在人脸编辑上实现LPIPS=0.001、SSIM=0.992,比基线提升7.8%~12.9%。
  • 适用于多种模型,显著增强身份保真度与视觉细节,适合高要求图像编辑场景。

矩形流模型中的图像编辑仍面临重建保真度与编辑灵活性之间的根本权衡。基于反演的方法存在轨迹偏移问题,而近期无反演方法如FlowEdit虽提供直接编辑路径,但仍需额外引导以提升结构保持能力。本文证明最优传输理论可为矩形流编辑的两种范式提供统一框架。我们提出一种零样本传输引导反演框架,在逆向扩散过程中利用最优传输;同时将最优传输原则扩展至无反演方法,通过传输优化速度场修正。结合传输引导可有效平衡不同矩形流编辑方法的重建准确率与编辑可控性。对于基于反演的编辑,本方法在人脸编辑基准上达到LPIPS=0.001、SSIM=0.992,较RF-Inversion在LSUN数据集上提升7.8%~12.9%。对于FlowEdit在FLUX和Stable Diffusion 3上的无反演编辑,各类编辑场景下均实现语义一致性与结构保持的持续改进。语义人脸编辑实验显示身份保真度提升11.2%,感知质量更优。统一的最优传输框架在两种编辑范式下均生成视觉逼真、细节保留出色的编辑结果。代码已开源:https://github.com/marianlupascu/OT-RF

原文摘要 · Abstract (English)

Image editing in rectified flow models remains challenging due to the fundamental trade-off between reconstruction fidelity and editing flexibility. While inversion-based methods suffer from trajectory deviation, recent inversion-free approaches like FlowEdit offer direct editing pathways but can benefit from additional guidance to improve structure preservation. In this work, we demonstrate that optimal transport theory provides a unified framework for improving both paradigms in rectified flow editing. We introduce a zero-shot transport-guided inversion framework that leverages optimal transport during the reverse diffusion process, and extend optimal transport principles to enhance inversion-free methods through transport-optimized velocity field corrections. Incorporating transport-based guidance can effectively balance reconstruction accuracy and editing controllability across different rectified flow editing approaches. For inversion-based editing, our method achieves high-fidelity reconstruction with LPIPS scores of 0.001 and SSIM of 0.992 on face editing benchmarks, observing 7.8% to 12.9% improvements over RF-Inversion on LSUN datasets. For inversion-free editing with FlowEdit on FLUX and Stable Diffusion 3, we demonstrate consistent improvements in semantic consistency and structure preservation across diverse editing scenarios. Our semantic face editing experiments show an 11.2% improvement in identity preservation and enhanced perceptual quality. The unified optimal transport framework produces visually compelling edits with superior detail preservation across both inversion-based and direct editing paradigms. Code is available for RF-Inversion and FlowEdit at: https://github.com/marianlupascu/OT-RF

图像编辑最优传输扩散模型结构保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。