融合传播与生成,让视频外扩更连贯真实
Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting

- 用光流传播+细调网络,保持运动一致性
- 参考引导的潜在传播,提升跨帧内容延续性
- 适合动态场景大范围视频外扩任务
视频外扩旨在扩展原始画面边界之外的内容,同时保持空间一致性和帧间时间连贯性。现有方法主要依赖大规模生成模型(如扩散模型),但存在隐式时间建模和空间上下文有限的问题,导致帧内与帧间不一致,尤其在动态场景和大范围外扩时更为明显。为此,我们提出Seen-to-Scene框架,融合传播与生成范式。该方法利用预训练于视频修复的光流完成网络,经端到端微调以弥合领域差距,并重建连贯运动场;进一步引入参考引导的潜在传播机制,有效传递源内容。大量实验表明,本方法在推理效率与视觉真实感方面均优于需输入定制化的现有最先进方法,实现更高时间连贯性。
原文摘要 · Abstract (English)
Video outpainting aims to expand the visible content of a video beyond the original frame boundaries while preserving spatial fidelity and temporal coherence across frames. Existing methods primarily rely on large-scale generative models, such as diffusion models. However, generationbased approaches suffer from implicit temporal modeling and limited spatial context. These limitations lead to intraframe and inter-frame inconsistencies, which become particularly pronounced in dynamic scenes and large outpainting scenarios. To overcome these challenges, we propose Seen-to-Scene, a novel framework that unifies propagationbased and generation-based paradigms for video outpainting. Specifically, Seen-to-Scene leverages flow-based propagation with a flow completion network pre-trained for video inpainting, which is fine-tuned in an end-to-end manner to bridge the domain gap and reconstruct coherent motion fields. To further improve the efficiency and reliability of propagation, we introduce a reference-guided latent propagation that effectively propagates source content across frames. Extensive experiments demonstrate that our method achieves superior temporal coherence and visual realism with efficient inference, surpassing even prior state-of-the-art methods that require input-specific adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。