Any4D实现高效统一的4D场景重建,支持多模态输入并显著提升精度与速度。
Any4D: Unified Feed-Forward Metric 4D Reconstruction
- 采用模块化表示,分别编码局部视角与全局世界坐标下的运动和几何信息
- 在多种场景下误差降低2-3倍,推理速度比现有方法快15倍
- 适合需要高精度实时4D重建的自动驾驶与机器人应用
我们提出Any4D,一种可扩展的多视角变换器,用于度量尺度、稠密的前馈式4D重建。与以往仅关注双视图稠密场景流或稀疏3D点跟踪的方法不同,Any4D直接为N帧生成每像素的运动与几何预测。此外,不同于其他单目RGB视频的4D重建方法,Any4D可在有可用时处理额外模态,如RGB-D图像、基于IMU的自运动信息及雷达多普勒测量。其核心创新在于对4D场景的模块化表示:各视角的4D预测通过本地相机坐标系中的自心因子(深度图与相机内参)和全局世界坐标系中的他心因子(相机外参与场景流)进行编码。我们在多种设置下均取得更优性能——精度提升2-3倍,计算效率提高15倍,为多个下游应用开辟了新路径。
原文摘要 · Abstract (English)
We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on either 2-view dense scene flow or sparse 3D point tracking. Moreover, unlike other recent methods for 4D reconstruction from monocular RGB videos, Any4D can process additional modalities and sensors such as RGB-D frames, IMU-based egomotion, and Radar Doppler measurements, when available. One of the key innovations that allows for such a flexible framework is a modular representation of a 4D scene; specifically, per-view 4D predictions are encoded using a variety of egocentric factors (depthmaps and camera intrinsics) represented in local camera coordinates, and allocentric factors (camera extrinsics and scene flow) represented in global world coordinates. We achieve superior performance across diverse setups - both in terms of accuracy (2-3X lower error) and compute efficiency (15X faster), opening avenues for multiple downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。