无需复杂计算,测试时引导即可高效实现图像视频编辑
When Test-Time Guidance Is Enough: Fast Image and Video Editing with Diffusion Guidance
- 直接使用测试时引导,跳过耗时的向量-雅可比乘积计算
- 在大规模图像视频基准上表现媲美甚至超越训练型方法
- 适合追求快速部署与轻量级编辑的应用场景
文本驱动的图像和视频编辑可自然建模为修复问题,即对遮蔽区域进行重建以同时保持与已有内容和编辑提示的一致性。近期基于扩散模型和流模型的测试时引导方法为此任务提供了理论框架;然而,现有方法依赖于昂贵的向量-雅可比乘积(VJP)计算来近似难以处理的引导项,限制了其实际应用。本文基于Moufad等(2025)的工作,提供对其无VJP近似的理论分析,并大幅扩展其在大规模图像和视频编辑基准上的实证评估。结果表明,仅使用测试时引导即可达到与训练型方法相当,甚至在某些情况下更优的性能。
原文摘要 · Abstract (English)
Text-driven image and video editing can be naturally cast as inpainting problems, where masked regions are reconstructed to remain consistent with both the observed content and the editing prompt. Recent advances in test-time guidance for diffusion and flow models provide a principled framework for this task; however, existing methods rely on costly vector--Jacobian product (VJP) computations to approximate the intractable guidance term, limiting their practical applicability. Building upon the recent work of Moufad et al. (2025), we provide theoretical insights into their VJP-free approximation and substantially extend their empirical evaluation to large-scale image and video editing benchmarks. Our results demonstrate that test-time guidance alone can achieve performance comparable to, and in some cases surpass, training-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。