无需标注数据,实现佩戴式视频的快速4D场景重建。
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
- 基于自监督学习,统一估计相机内参、位姿和深度。
- 在无标签数据上实现优于所有基线的稠密点云序列重建。
- 适合关注可穿戴设备场景理解的研究者与开发者。
佩戴式视频为理解人类与物理世界的交互提供了宝贵视角,但其几何与动态的稠密场景重建仍面临挑战。由于缺乏高质量标注数据,现有监督学习方法效果受限。本文提出EgoMono4D,首个面向无标签佩戴式视频的自监督动态4D重建模型。该模型在预训练单帧深度与内参估计基础上,引入相机位姿估计,并在大规模未标注佩戴式视频上对齐多帧结果,构建端到端的快速前向框架。在域内与零样本泛化设置下均优于所有基线,实现高效、稠密、可泛化的点云序列重建。代码与模型已开源。
原文摘要 · Abstract (English)
Egocentric videos provide valuable insights into human interactions with the physical world, which has sparked growing interest in the computer vision and robotics communities. A critical challenge in fully understanding the geometry and dynamics of egocentric videos is dense scene reconstruction. However, the lack of high-quality labeled datasets in this field has hindered the effectiveness of current supervised learning methods. In this work, we aim to address this issue by exploring an self-supervised dynamic scene reconstruction approach. We introduce EgoMono4D, a novel model that unifies the estimation of multiple variables necessary for Egocentric Monocular 4D reconstruction, including camera intrinsic, camera poses, and video depth, all within a fast feed-forward framework. Starting from pretrained single-frame depth and intrinsic estimation model, we extend it with camera poses estimation and align multi-frame results on large-scale unlabeled egocentric videos. We evaluate EgoMono4D in both in-domain and zero-shot generalization settings, achieving superior performance in dense pointclouds sequence reconstruction compared to all baselines. EgoMono4D represents the first attempt to apply self-supervised learning for pointclouds sequence reconstruction to the label-scarce egocentric field, enabling fast, dense, and generalizable reconstruction. The interactable visualization, code and trained models are released https://egomono4d.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。