arXiv:2506.05011cs.CV2025-06AAAI被引 3

用高斯溅射实现无人机视角下多人动态场景的逼真渲染

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting

  • 结合3D基础模型与人体网格重建,从单目视频恢复动态场景
  • 通过人体-场景接触点解决尺度模糊,实现世界坐标对齐
  • 在三个无人机数据集上提升1.5dB PSNR,适合真实飞行场景应用

尽管动态神经渲染取得显著进展,现有方法未能解决无人机拍摄场景带来的独特挑战,尤其是单目相机、俯视视角及多个小而移动的人体,这些在现有数据集中缺乏充分表征。本文提出UAV4D框架,实现无人机捕获动态真实场景的逼真渲染。具体而言,我们仅利用单目视频数据,无需额外传感器即可重建包含多行人动态的场景。通过融合3D基础模型与人体网格重建模型,恢复场景背景与人体。提出新方法通过识别人体-场景接触点解决场景尺度模糊问题,将人体与场景统一至世界坐标系。同时,利用SMPL模型与背景网格初始化高斯溅射,实现整体场景渲染。我们在三个复杂无人机数据集:VisDrone、Manipal-UAV和Okutama-Action上评估,每个数据集含10~50名行人。结果表明,相比现有方法,本方法在新视角合成上实现1.5 dB PSNR提升,视觉清晰度更优。

原文摘要 · Abstract (English)

Despite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple small, moving humans, which are not adequately represented in existing datasets. In this work, we introduce UAV4D, a framework for enabling photorealistic rendering for dynamic real-world scenes captured by UAVs. Specifically, we address the challenge of reconstructing dynamic scenes with multiple moving pedestrians from monocular video data without the need for additional sensors. We use a combination of a 3D foundation model and a human mesh reconstruction model to reconstruct both the scene background and humans. We propose a novel approach to resolve the scene scale ambiguity and place both humans and the scene in world coordinates by identifying human-scene contact points. Additionally, we exploit the SMPL model and background mesh to initialize Gaussian splats, enabling holistic scene rendering. We evaluated our method on three complex UAV-captured datasets: VisDrone, Manipal-UAV, and Okutama-Action, each with distinct characteristics and 10~50 humans. Our results demonstrate the benefits of our approach over existing methods in novel view synthesis, achieving a 1.5 dB PSNR improvement and superior visual sharpness.

神经渲染无人机影像高斯溅射人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。