高分辨率长视频外扩生成,用分阶段策略实现稳定一致的视觉拓展。
HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos

- 先建全局粗略引导图,融合长时结构与短时动态信息。
- 在粗引导下进行高分辨率生成,支持大范围空间扩展和长序列。
- 适合需要视频画面外延的影视、广告等场景使用。
视频外扩旨在生成超出原始视频空间范围的合理视觉内容,在适配多样显示格式中起关键作用。为支持此类应用,需在长序列上实现大范围空间外推。然而,现有方法大多仅解决单一挑战,或缺乏确保全局时空一致性的明确机制,导致明显局限。本文提出HL-OutPaint,一种面向长序列高分辨率视频外扩的框架。该方法采用分阶段粗到精策略,构建全局粗略引导(GCG),以低分辨率表示捕捉视频整体结构与主导运动。不同于简单下采样,GCG通过新颖的全局-局部帧交换机制,将稀疏全局关键帧与局部时间窗口结合,在采样过程中实现信息交互,从而统一编码长期结构一致性与短期时序动态。基于此表示,系统执行高分辨率外扩生成,输出空间细节丰富且时序一致的内容。通过分离全局结构建模与精细合成,本框架在大范围空间拓展与长视频序列下均实现稳定、连贯生成。大量实验表明,其在涉及宽幅空间外推与长视频序列的挑战性场景中优于现有方法。
原文摘要 · Abstract (English)
Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role in adapting videos to diverse display formats. To support such use cases, it must enable large spatial extrapolation over long sequences. However, most existing methods address only one of these challenges or lack explicit mechanisms for ensuring global spatio-temporal consistency, leading to notable limitations. In this paper, we propose HL-OutPaint, a high-resolution video outpainting framework for long sequences. Our approach follows a coarse-to-fine strategy with a two-stage pipeline. We first construct Global Coarse Guidance (GCG), a low-resolution representation that captures global structure and dominant motion across the video. Unlike naive downsampling, GCG is built via a novel global-local frame swapping mechanism that couples sparse global keyframes with local temporal windows and exchanges information during sampling. This enables GCG to encode both long-term structural consistency and short-term temporal dynamics in a unified representation. Guided by this representation, HL-OutPaint then performs high-resolution outpainting to generate spatially detailed and temporally consistent content. By separating global structure modeling from fine-grained synthesis, our framework achieves stable, coherent generation for large spatial expansion and long video sequences. Extensive experiments show that HL-OutPaint outperforms existing methods in challenging scenarios involving wide spatial extrapolation and long video sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。