arXiv:2504.03041cs.CV2025-04被引 1

无需提示词的视频修复框架,精准移除真实场景中的人与物品

VIP: Video Inpainting Pipeline for Real World Human Removal

  • 基于文本到视频模型改进,加入运动模块与变分自编码器逐步去噪
  • 在复杂场景下实现高时序一致性与视觉保真度,超越现有方法
  • 适合需要长视频高分辨率修复的应用,如影视剪辑与安防监控

真实世界高分辨率视频中的人物与行人移除面临巨大挑战,尤其在保证高质量输出、时间一致性及处理涉及人体、随身物品和影子的复杂交互方面。本文提出VIP(Video Inpainting Pipeline),一种用于真实场景人物移除的新型无提示词视频修复框架。VIP通过引入运动模块增强先进文本到视频模型,并采用变分自编码器(VAE)在潜在空间进行渐进式去噪。此外,我们设计了高效的人体与随身物品分割方法以生成精确掩码。大量实验表明,VIP在多样真实场景中实现了优异的时间一致性与视觉保真度,在具有挑战性的数据集上优于现有最先进方法。主要贡献包括构建了VIP流水线、参考帧融合技术以及双融合潜在分割精炼方法,有效应对长视频高分辨率修复中的复杂性。

原文摘要 · Abstract (English)

Inpainting for real-world human and pedestrian removal in high-resolution video clips presents significant challenges, particularly in achieving high-quality outcomes, ensuring temporal consistency, and managing complex object interactions that involve humans, their belongings, and their shadows. In this paper, we introduce VIP (Video Inpainting Pipeline), a novel promptless video inpainting framework for real-world human removal applications. VIP enhances a state-of-the-art text-to-video model with a motion module and employs a Variational Autoencoder (VAE) for progressive denoising in the latent space. Additionally, we implement an efficient human-and-belongings segmentation for precise mask generation. Sufficient experimental results demonstrate that VIP achieves superior temporal consistency and visual fidelity across diverse real-world scenarios, surpassing state-of-the-art methods on challenging datasets. Our key contributions include the development of the VIP pipeline, a reference frame integration technique, and the Dual-Fusion Latent Segment Refinement method, all of which address the complexities of inpainting in long, high-resolution video sequences.

视频修复人体移除时序一致性潜在空间去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。