arXiv:2512.03905cs.CV2025-12

通过时空对应关系提升零样本视频生成一致性

Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence

  • 融合帧内与帧间对应关系构建更强时空约束
  • 生成视频在语义和时间上保持高度一致
  • 适合需要高质量视频编辑的科研与创作人群

文本到图像扩散模型的显著成功激发了其在视频应用中的广泛探索。零样本技术旨在不进行额外训练的情况下将图像扩散模型应用于视频。现有方法多聚焦于将帧间对应关系融入注意力机制,但其施加的软约束不足以有效识别需关注的特征,可能导致时间不一致。本文提出FRESCO,通过结合帧内与帧间对应关系,构建更鲁棒的空间-时间约束,确保帧间语义相似内容的一致变换。该方法超越注意力引导,显式优化特征表示,显著提升输入视频的时空一致性,增强操作后视频的视觉连贯性。我们在两个零样本任务——视频到视频转换与文本引导视频编辑上验证了FRESCO的有效性。大量实验表明,该框架能生成高质量、连贯的视频,在当前零样本方法中实现显著进步。

原文摘要 · Abstract (English)

The remarkable success in text-to-image diffusion models has motivated extensive investigation of their potential for video applications. Zero-shot techniques aim to adapt image diffusion models for videos without requiring further model training. Recent methods largely emphasize integrating inter-frame correspondence into attention mechanisms. However, the soft constraint applied to identify the valid features to attend is insufficient, which could lead to temporal inconsistency. In this paper, we present FRESCO, which integrates intra-frame correspondence with inter-frame correspondence to formulate a more robust spatial-temporal constraint. This enhancement ensures a consistent transformation of semantically similar content between frames. Our method goes beyond attention guidance to explicitly optimize features, achieving high spatial-temporal consistency with the input video, significantly enhancing the visual coherence of manipulated videos. We verify FRESCO adaptations on two zero-shot tasks of video-to-video translation and text-guided video editing. Comprehensive experiments demonstrate the effectiveness of our framework in generating high-quality, coherent videos, highlighting a significant advance over current zero-shot methods.

视频生成扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。