视频修复新框架,支持任意长度修复与编辑,效果更自然。
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
- 双流架构分离背景建模与内容生成,提升修复精度。
- 支持任意长度视频修复,单个模型覆盖全时长场景。
- 构建超大规模数据集,推动视频修复技术发展。
视频修复旨在恢复受损内容,现有方法在生成完全遮挡物体或平衡背景保留与前景生成方面存在局限。为此,我们提出 VideoPainter,一种新型双流架构,通过仅占主干6%参数的高效上下文编码器处理遮挡视频,并将背景上下文提示注入预训练视频DiT,实现即插即用的语义一致内容生成。该架构显著降低学习复杂度,支持细粒度背景融合。我们还引入目标区域ID重采样技术,实现任意长度视频修复,极大提升实用性。此外,构建基于视觉理解模型的可扩展数据集流水线,贡献了目前最大规模的视频修复数据集与基准(VPData and VPBench),包含超过39万条多样化视频片段。以修复为基线,进一步探索视频编辑及成对编辑数据生成任务,展现出优异性能和显著应用潜力。大量实验表明,VideoPainter在八项关键指标上表现卓越,涵盖视频质量、掩码区域保持与文本一致性。
原文摘要 · Abstract (English)
Video inpainting, which aims to restore corrupted video content, has experienced substantial progress. Despite these advances, existing methods, whether propagating unmasked region pixels through optical flow and receptive field priors, or extending image-inpainting models temporally, face challenges in generating fully masked objects or balancing the competing objectives of background context preservation and foreground generation in one model, respectively. To address these limitations, we propose a novel dual-stream paradigm VideoPainter that incorporates an efficient context encoder (comprising only 6% of the backbone parameters) to process masked videos and inject backbone-aware background contextual cues to any pre-trained video DiT, producing semantically consistent content in a plug-and-play manner. This architectural separation significantly reduces the model's learning complexity while enabling nuanced integration of crucial background context. We also introduce a novel target region ID resampling technique that enables any-length video inpainting, greatly enhancing our practical applicability. Additionally, we establish a scalable dataset pipeline leveraging current vision understanding models, contributing VPData and VPBench to facilitate segmentation-based inpainting training and assessment, the largest video inpainting dataset and benchmark to date with over 390K diverse clips. Using inpainting as a pipeline basis, we also explore downstream applications including video editing and video editing pair data generation, demonstrating competitive performance and significant practical potential. Extensive experiments demonstrate VideoPainter's superior performance in both any-length video inpainting and editing, across eight key metrics, including video quality, mask region preservation, and textual coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。