评估动态3D高斯溅射在第一人称场景重建中的表现,发现静态内容重建更差。
Bringing a Personal Point of View: Evaluating Dynamic 3D Gaussian Splatting for Egocentric Scene Reconstruction
- 用配对的第一人称与第三人称视频对比测试动态3DGS模型
- 第一人称视角下峰值信噪比更低,主要因静态内容重建不佳
- 提示需专门设计第一人称场景重建方法,建议分区域评估
第一人称视频为理解人类感知与交互提供了独特视角,在增强现实、机器人和辅助技术中日益重要。然而,快速相机运动和复杂场景动态给从该视角进行3D重建带来了重大挑战。尽管3D高斯溅射(3DGS)已成为高效高质量新视图合成的前沿方法,但针对单目视频动态场景重建的变体在第一人称视频上的评估极为少见。现有模型能否泛化到此场景,或是否需要专为第一人称设计的方案,仍不明确。本文使用来自EgoExo4D数据集的配对第一人称-第三人称视频,评估了动态单目3DGS模型在两类视频上的表现。结果表明,第一人称视角下的重建质量始终较低。分析显示,这一差异主要源于静态内容的重建问题,而非动态部分。研究强调了当前方法的局限性,并推动开发专用于第一人称场景的重建方案,同时指出应分别评估视频中的静态与动态区域。
原文摘要 · Abstract (English)
Egocentric video provides a unique view into human perception and interaction, with growing relevance for augmented reality, robotics, and assistive technologies. However, rapid camera motion and complex scene dynamics pose major challenges for 3D reconstruction from this perspective. While 3D Gaussian Splatting (3DGS) has become a state-of-the-art method for efficient, high-quality novel view synthesis, variants, that focus on reconstructing dynamic scenes from monocular video are rarely evaluated on egocentric video. It remains unclear whether existing models generalize to this setting or if egocentric-specific solutions are needed. In this work, we evaluate dynamic monocular 3DGS models on egocentric and exocentric video using paired ego-exo recordings from the EgoExo4D dataset. We find that reconstruction quality is consistently lower in egocentric views. Analysis reveals that the difference in reconstruction quality, measured in peak signal-to-noise ratio, stems from the reconstruction of static, not dynamic, content. Our findings underscore current limitations and motivate the development of egocentric-specific approaches, while also highlighting the value of separately evaluating static and dynamic regions of a video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。