arXiv:2509.18566cs.CVcs.RO2025-09被引 2

用事件相机提升单目视频中动态人体与场景的重建质量

Event-guided 3D Gaussian Splatting for Dynamic Human and Scene Reconstruction

  • 用3D高斯点云统一建模人体与场景,仅人体点云可变形
  • 通过事件流匹配渲染亮度变化,有效缓解高速运动中的模糊
  • 无需外部人体掩码,适合快速运动场景重建

从单目视频中同时重建动态人体与静态场景仍具挑战,尤其在快速运动时,RGB帧易出现运动模糊。事件相机具有微秒级时间分辨率,更适用于动态人体重建。为此,我们提出一种新型事件引导的人体-场景重建框架,通过3D高斯点云联合建模人体与场景。统一的3D高斯集合包含可学习语义属性:仅被分类为人体的高斯点云进行形变以实现动画,场景点云保持静态。为缓解模糊问题,我们设计了一种事件引导损失函数,将连续渲染间的模拟亮度变化与事件流进行匹配,显著提升高速运动区域的局部保真度。该方法无需外部人体掩码,也无需分别管理多个高斯集合。在两个基准数据集ZJU-MoCap-Blur和MMHPSD-Blur上,本方法达到当前最优性能,相较强基线在PSNR/SSIM上有明显提升,且LPIPS降低,尤其在高速主体上表现突出。

原文摘要 · Abstract (English)

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cameras exhibit distinct advantages, e.g., microsecond temporal resolution, making them a superior sensing choice for dynamic human reconstruction. Accordingly, we present a novel event-guided human-scene reconstruction framework that jointly models human and scene from a single monocular event camera via 3D Gaussian Splatting. Specifically, a unified set of 3D Gaussians carries a learnable semantic attribute; only Gaussians classified as human undergo deformation for animation, while scene Gaussians stay static. To combat blur, we propose an event-guided loss that matches simulated brightness changes between consecutive renderings with the event stream, improving local fidelity in fast-moving regions. Our approach removes the need for external human masks and simplifies managing separate Gaussian sets. On two benchmark datasets, ZJU-MoCap-Blur and MMHPSD-Blur, it delivers state-of-the-art human-scene reconstruction, with notable gains over strong baselines in PSNR/SSIM and reduced LPIPS, especially for high-speed subjects.

3D重建事件相机动态人体高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。