arXiv:2503.23035cs.CV2025-03NeurIPS被引 6

FreeInv通过随机变换隐空间实现高效图像视频逆向生成。

FreeInv: Free Lunch for Improving DDIM Inversion

  • 随机变换隐变量并保持重建与逆向一致,降低轨迹偏差。
  • 在PIE和DAVIS数据集上显著优于传统DDIM逆向,计算开销低。
  • 适合需要高保真度的图像视频编辑任务,可无缝集成现有方法。

传统的DDIM逆向过程常因重构时的隐空间轨迹偏离原反向轨迹而产生偏差。以往方法或学习缓解偏差,或设计复杂补偿策略,但计算成本高。本文提出一种近乎零成本的方法FreeInv:在逆向与重构的对应时间步中对隐表示进行相同随机变换。从统计角度看,多个轨迹的集成能期望降低轨迹不匹配误差。理论分析与实验证明,FreeInv实现了高效的多轨迹集成。该方法可自由嵌入现有基于逆向的图像与视频编辑技术中。尤其在视频序列逆向时,显著提升保真度与效率。在PIE基准和DAVIS数据集上的定量与定性评估显示,FreeInv明显优于传统DDIM逆向,且性能媲美当前最优方法,同时具备更优的计算效率。

原文摘要 · Abstract (English)

Naive DDIM inversion process usually suffers from a trajectory deviation issue, i.e., the latent trajectory during reconstruction deviates from the one during inversion. To alleviate this issue, previous methods either learn to mitigate the deviation or design cumbersome compensation strategy to reduce the mismatch error, exhibiting substantial time and computation cost. In this work, we present a nearly free-lunch method (named FreeInv) to address the issue more effectively and efficiently. In FreeInv, we randomly transform the latent representation and keep the transformation the same between the corresponding inversion and reconstruction time-step. It is motivated from a statistical perspective that an ensemble of DDIM inversion processes for multiple trajectories yields a smaller trajectory mismatch error on expectation. Moreover, through theoretical analysis and empirical study, we show that FreeInv performs an efficient ensemble of multiple trajectories. FreeInv can be freely integrated into existing inversion-based image and video editing techniques. Especially for inverting video sequences, it brings more significant fidelity and efficiency improvements. Comprehensive quantitative and qualitative evaluation on PIE benchmark and DAVIS dataset shows that FreeInv remarkably outperforms conventional DDIM inversion, and is competitive among previous state-of-the-art inversion methods, with superior computation efficiency.

图像逆向扩散模型视频生成效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。