让3D场景重建在时间上连续,用事件数据填补帧间空白
Interp3R: Continuous-time 3D Geometry Estimation with Frames and Events
- 用事件数据插值帧式模型生成的点云,实现任意时刻的几何估计
- 在合成与真实数据上均超越现有两阶段方法,精度显著提升
- 首次实现基于点云的3D模型时间连续化,适合动态场景应用
近年来,以点云为基础的3D视觉基础模型(如DUSt3R)受到广泛关注,其在多种场景下展现出优异的精度与强泛化能力。然而,这些方法仅能恢复图像捕获瞬间的场景几何,无法刻画帧间盲区中的场景演化。本文提出Interp3R,据我们所知首个将帧式点云模型扩展为任意时间点几何估计的方法。该方法利用异步事件数据对帧式模型输出的点云进行插值,实现时空连续的几何表示。通过将插值后的点云与原始帧预测点云对齐,联合恢复深度和相机位姿。模型仅在合成数据上训练,却在大量合成与真实世界基准测试中表现出强泛化能力。大量实验表明,Interp3R显著优于现有两阶段方法——先进行2D视频帧插值,再进行3D几何估计。
原文摘要 · Abstract (English)
In recent years, 3D visual foundation models pioneered by pointmap-based approaches such as DUSt3R have attracted a lot of interest, achieving impressive accuracy and strong generalization across diverse scenes. However, these methods are inherently limited to recovering scene geometry only at the discrete time instants when images are captured, leaving the scene evolution during the blind time between consecutive frames largely unexplored. We introduce Interp3R, to the best of our knowledge the first method that enhances pointmap-based models to estimate depth and camera poses at arbitrary time instants. Interp3R leverages asynchronous event data to interpolate pointmaps produced by frame-based models, enabling temporally continuous geometric representations. Depth and camera poses are then jointly recovered by aligning the interpolated pointmaps together with those predicted by the underlying frame-based models into a consistent spatial framework. We train Interp3R exclusively on a synthetic dataset, yet demonstrate strong generalization across a wide range of synthetic and real-world benchmarks. Extensive experiments show that Interp3R outperforms by a considerable margin state-of-the-art baselines that follow a two-stage pipeline of 2D video frame interpolation followed by 3D geometry estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。