arXiv:2603.29036cs.CV2026-03

用合成数据训练模型,从第一视角步行视频中干净移除行人及阴影。

Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos

  • 构建半合成视频对数据集,包含背景与叠加行人的模拟片段。
  • 微调Casper扩散模型,在复杂背景中成功移除多人且保留真实感。
  • 生成的无行人视频可有效支持城市三维/四维场景重建,适合视觉建模者。

第一视角步行视频为全球环境的视觉建模提供了丰富的图像数据,但因人群密集和摄像头高度接近人眼,画面中常出现大量行人,影响建模效果。本文提出一种生成式算法,可真实地移除视频中的行人及其阴影。核心是构建一个半合成视频对数据集,每对包含一个仅含环境的背景片段,以及在该背景上叠加模拟行人的复合片段。前景与背景均来自世界各地的真实步行视频,确保视觉多样性。利用该数据集微调当前最先进的Casper视频扩散模型,用于物体与光影修复。结果表明,改进后的模型在定性和定量评估中均显著优于原始Casper,在高人流量、复杂背景的视频中表现更佳。最终,生成的无行人视频可用于成功构建城市区域的3D/4D模型。

原文摘要 · Abstract (English)

Egocentric "walking tour" videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence of humans in frames of these videos due to crowds and eye-level camera perspectives mitigates their usefulness in environment modeling applications. We focus on addressing this challenge by developing a generative algorithm that can realistically remove (i.e., inpaint) humans and their associated shadow effects from walking tour videos. Key to our approach is the construction of a rich semi-synthetic dataset of video clip pairs to train this generative model. Each pair in the dataset consists of an environment-only background clip, and a composite clip of walking humans with simulated shadows overlaid on the background. We randomly sourced both foreground and background components from real egocentric walking tour videos around the world to maintain visual diversity. We then used this dataset to fine-tune the state-of-the-art Casper video diffusion model for object and effects inpainting, and demonstrate that the resulting model performs far better than Casper both qualitatively and quantitatively at removing humans from walking tour clips with significant human presence and complex backgrounds. Finally, we show that the resulting generated clips can be used to build successful 3D/4D models of urban locations.

视频修复扩散模型场景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。