arXiv:2510.22970cs.CV2025-10

无需训练即可实现视频编辑的时序一致性,自动选关键帧压缩成语义锚点。

VALA: Learning Latent Anchors for Training-Free and Temporally Consistent

  • 用变分对齐模块自动生成关键帧并压缩为语义锚点
  • 在真实视频数据集上达到最佳时序一致性和编辑质量
  • 适合追求高效、免训练视频编辑的研究者和开发者

近期训练自由视频编辑方法利用预训练文本到图像扩散模型,实现了轻量级且精准的跨帧生成。然而,现有方法通常依赖启发式帧选择来维持DDIM反演过程中的时序一致性,引入人工偏倚并降低端到端推理的可扩展性。本文提出VALA(Variational Alignment for Latent Anchors),一种变分对齐模块,可自适应选择关键帧,并将其潜在特征压缩为语义锚点以保证视频编辑的一致性。为学习有意义的分配,VALA采用具有对比学习目标的变分框架,从而将跨帧潜在表示转换为保留内容与时序连贯性的压缩潜在锚点。该方法可完全集成至基于训练自由文本到图像模型的视频编辑系统中。在真实世界视频编辑基准上的大量实验表明,VALA在反演保真度、编辑质量和时序一致性方面均达到当前最优性能,同时相比先前方法具有更高效率。

原文摘要 · Abstract (English)

Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to maintain temporal consistency during DDIM inversion, which introduces manual bias and reduces the scalability of end-to-end inference. In this paper, we propose~\textbf{VALA} (\textbf{V}ariational \textbf{A}lignment for \textbf{L}atent \textbf{A}nchors), a variational alignment module that adaptively selects key frames and compresses their latent features into semantic anchors for consistent video editing. To learn meaningful assignments, VALA propose a variational framework with a contrastive learning objective. Therefore, it can transform cross-frame latent representations into compressed latent anchors that preserve both content and temporal coherence. Our method can be fully integrated into training-free text-to-image based video editing models. Extensive experiments on real-world video editing benchmarks show that VALA achieves state-of-the-art performance in inversion fidelity, editing quality, and temporal consistency, while offering improved efficiency over prior methods.

视频编辑扩散模型时序一致免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。