arXiv:2506.22868cs.CVcs.AI2025-06

无需训练即可实现视频编辑,保持时空一致性。

STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing

  • 用新型时空相关性评分引导隐空间优化,避免3D注意力开销。
  • 在显著域变换下仍保持视觉真实性和时间连贯性。
  • 适合需要快速编辑且不依赖训练数据的场景。

以往文本引导的视频编辑方法常存在时间不一致、运动失真及领域转换能力有限等问题。我们归因于编辑过程中对时空像素相关性建模不足。为此,提出STR-Match,一种无需训练的视频编辑算法,通过新颖的时空相关性评分(STR score)指导隐空间优化,在不使用计算昂贵的3D注意力机制的前提下,利用2D空间注意力与1D时间模块捕捉相邻帧间的时空像素相关性。结合隐空间掩码,该方法生成视觉自然、时空一致的视频,在大幅域变换下仍能保留源视频关键视觉属性。大量实验表明,STR-Match在视觉质量与时空一致性上均持续优于现有方法。

原文摘要 · Abstract (English)

Previous text-guided video editing methods often suffer from temporal inconsistency, motion distortion, and-most notably-limited domain transformation. We attribute these limitations to insufficient modeling of spatiotemporal pixel relevance during the editing process. To address this, we propose STR-Match, a training-free video editing algorithm that produces visually appealing and spatiotemporally coherent videos through latent optimization guided by our novel STR score. The score captures spatiotemporal pixel relevance across adjacent frames by leveraging 2D spatial attention and 1D temporal modules in text-to-video (T2V) diffusion models, without the overhead of computationally expensive 3D attention mechanisms. Integrated into a latent optimization framework with a latent mask, STR-Match generates temporally consistent and visually faithful videos, maintaining strong performance even under significant domain transformations while preserving key visual attributes of the source. Extensive experiments demonstrate that STR-Match consistently outperforms existing methods in both visual quality and spatiotemporal consistency.

视频编辑扩散模型时空一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。