用重叠共去噪实现长视频任意长度修复与扩展,无拼接痕迹。
Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising
- 通过高阶求解器和重叠融合策略保持长视频时序一致性。
- 在数百帧上修复/添加物体,PSNR/SSIM优于基线模型。
- 仅用LoRA微调大模型,参数高效且生成无伪影。
长视频生成仍是核心挑战,而视频修复与扩展的高可控性尤为困难。为同时解决这两项难题,我们提出一种统一的长视频修复与扩展方法,将文本到视频扩散模型拓展为可生成任意长度、空间编辑的高保真视频。该方法利用LoRA高效微调阿里万像2.1等预训练视频扩散模型以合成遮蔽区域内容,并采用重叠-融合的时间共去噪策略结合高阶求解器,确保长序列的一致性。相比以往工作在固定长度片段或存在拼接伪影的局限,本系统支持任意长度生成与编辑,且无明显接缝或漂移。我们在复杂修复/扩展任务中验证了该方法,涵盖数百帧上的对象编辑或新增,其质量(PSNR/SSIM)和感知真实度(LPIPS)均优于基线模型如万像2.1与VACE。该方法以极小开销实现实用的长距离视频编辑,在参数效率与性能间取得良好平衡。
原文摘要 · Abstract (English)
Generating long videos remains a fundamental challenge, and achieving high controllability in video inpainting and outpainting is particularly demanding. To address both of these challenges simultaneously and achieve controllable video inpainting and outpainting for long video clips, we introduce a novel and unified approach for long video inpainting and outpainting that extends text-to-video diffusion models to generate arbitrarily long, spatially edited videos with high fidelity. Our method leverages LoRA to efficiently fine-tune a large pre-trained video diffusion model like Alibaba's Wan 2.1 for masked region video synthesis, and employs an overlap-and-blend temporal co-denoising strategy with high-order solvers to maintain consistency across long sequences. In contrast to prior work that struggles with fixed-length clips or exhibits stitching artifacts, our system enables arbitrarily long video generation and editing without noticeable seams or drift. We validate our approach on challenging inpainting/outpainting tasks including editing or adding objects over hundreds of frames and demonstrate superior performance to baseline methods like Wan 2.1 model and VACE in terms of quality (PSNR/SSIM), and perceptual realism (LPIPS). Our method enables practical long-range video editing with minimal overhead, achieved a balance between parameter efficient and superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。