arXiv:2510.21461cs.CV2025-10

用帧间距先验提升视频修复的时序一致性与内容稳定性。

Enhancing Video Inpainting with Aligned Frame Interval Guidance

  • 将视频修复拆解为多帧图像修复和运动传播,引入帧间距先验引导生成。
  • 通过帧内容传播模块实现跨帧信息扩散,减少内容退化现象。
  • 适合需要高时序稳定性的视频修复任务,尤其对长视频修复有帮助。

基于图像到视频(I2V)的视频修复方法虽利用单图先验并建模跨掩码帧的时序一致性,但仍存在视频块内内容严重退化的问题。此外,缺乏稳健的帧对齐机制导致块内与块间时空不稳定,难以有效控制整个视频。为此,我们提出VidPivot框架,将视频修复分解为两个子任务:多帧一致的图像修复和掩码区域运动传播。该方法引入帧间距先验作为时空引导信号,设计了FrameProp模块,通过拼接机制将参考帧内容扩散至后续帧以增强跨帧一致性。同时,专门的上下文控制器将这些连贯帧先验编码至I2V生成主干中,作为软约束抑制生成过程中的内容失真。大量实验表明,VidPivot在多个基准上表现优异,并能良好泛化至不同视频修复场景。

原文摘要 · Abstract (English)

Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from severe content degradation within video chunks. Furthermore, the absence of a robust frame alignment scheme compromises intra-chunk and inter-chunk spatiotemporal stability, resulting in insufficient control over the entire video. To address these limitations, we propose VidPivot, a novel framework that decouples video inpainting into two sub-tasks: multi-frame consistent image inpainting and masked area motion propagation. Our approach introduces frame interval priors as spatiotemporal cues to guide the inpainting process. To enhance cross-frame coherence, we design a FrameProp Module that implements a frame content propagation strategy, diffusing reference frame content into subsequent frames via a splicing mechanism. Additionally, a dedicated context controller encodes these coherent frame priors into the I2V generative backbone, effectively serving as soft constrain to suppress content distortion during generation. Extensive evaluations demonstrate that VidPivot achieves competitive performance across diverse benchmarks and generalizes well to different video inpainting scenarios.

视频修复时序一致性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。