从网络立体视频中学习物体3D运动,无需标注即可生成高质量动态3D重建。
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
- 融合相机位姿、立体深度与时间跟踪结果,自动构建高质量4D动态场景
- 生成包含长期运动轨迹的全局一致、伪度量3D点云数据集
- 可用于训练模型泛化到真实复杂场景,适合视觉理解与机器人应用
从图像中理解动态3D场景对机器人和场景重建至关重要。然而,由于难以获取真实标注,直接监督3D运动恢复仍具挑战。本文提出一种从互联网立体广角视频中挖掘高质量4D重建的方法。系统融合相机位姿估计、立体深度估计与时间跟踪结果,生成高保真动态3D重建。利用该方法构建大规模数据集,包含全局一致、伪度量的3D点云及长期运动轨迹。我们以该数据训练DUSt3R变体,使其能从真实图像对中预测结构与3D运动,验证了数据在跨场景泛化上的有效性。项目页面与数据见:https://stereo4d.github.io
原文摘要 · Abstract (English)
Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due to the fundamental difficulty of obtaining ground truth annotations. We present a system for mining high-quality 4D reconstructions from internet stereoscopic, wide-angle videos. Our system fuses and filters the outputs of camera pose estimation, stereo depth estimation, and temporal tracking methods into high-quality dynamic 3D reconstructions. We use this method to generate large-scale data in the form of world-consistent, pseudo-metric 3D point clouds with long-term motion trajectories. We demonstrate the utility of this data by training a variant of DUSt3R to predict structure and 3D motion from real-world image pairs, showing that training on our reconstructed data enables generalization to diverse real-world scenes. Project page and data at: https://stereo4d.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。