提出动态高斯分解方法,实现第一人称4D场景中背景、手部、物体的精细分离。
DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction

- 用可学习类别概率动态分配高斯点到背景、手、物体分支
- 在多个指标上提升1.70dB PSNR,实现最佳分离效果
- 适合需要精细场景编辑的AR/VR与具身AI应用
第一人称视频对下一代4D场景重建至关重要,应用于AR/VR与具身AI。然而,由于复杂的自我运动、遮挡及手物交互,重建动态第一人称场景极具挑战。现有分解方法不适用,或假设固定视角,或将动态信息合并为单一前景。为此,我们提出DP-DeGauss,一种面向第一人称4D重建的动态概率高斯分解框架。该方法基于COLMAP先验初始化统一3D高斯集合,为每个高斯点增加可学习类别概率,并动态路由至专用形变分支以建模背景、手部或物体。采用类别特定掩码提升解耦效果,引入亮度与运动流控制以优化静态渲染和动态重建。大量实验表明,DP-DeGauss在平均PSNR上较基线提升+1.70dB,SSIM与LPIPS也显著改善。更重要的是,该框架首次实现背景、手部与物体组件的最优解耦,支持显式、细粒度分离,为更直观的自我视角场景理解与编辑铺平道路。
原文摘要 · Abstract (English)
Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic first-person scenes is challenging due to complex ego-motion, occlusions, and hand-object interactions. Existing decomposition methods are ill-suited, assuming fixed viewpoints or merging dynamics into a single foreground. To address these limitations, we introduce DP-DeGauss, a dynamic probabilistic Gaussian decomposition framework for egocentric 4D reconstruction. Our method initializes a unified 3D Gaussian set from COLMAP priors, augments each with a learnable category probability, and dynamically routes them into specialized deformation branches for background, hands, or object modeling. We employ category-specific masks for better disentanglement and introduce brightness and motion-flow control to improve static rendering and dynamic reconstruction. Extensive experiments show that DP-DeGauss outperforms baselines by +1.70dB in PSNR on average with SSIM and LPIPS gains. More importantly, our framework achieves the first and state-of-the-art disentanglement of background, hand, and object components, enabling explicit, fine-grained separation, paving the way for more intuitive ego scene understanding and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。