用事件相机实现高精度全局3D重建,突破传统方法局限。
Event3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature Aggregation

- 将事件流转为时空体素,通过时序注意力融合动态信息。
- 自监督预训练提升时间建模能力,仅需少量标注数据即可精调。
- 适合低光照、高速运动场景下的机器人三维感知任务。
鲁棒的3D重建对机器人和具身感知至关重要。尽管基于RGB图像的前馈方法(如DUSt3R)已在密集3D重建中取得显著进展,实现了全局几何一致性与强泛化能力,但将其扩展至事件相机仍面临挑战:事件相机具有异步、稀疏、高度动态的特性,且缺乏大规模标注数据集。本文提出Event3R,一种直接从异步事件流生成全局一致3D点云的前馈框架。Event3R将输入事件表示为时空体素,通过时序注意力模块实现时间感知特征融合,增强时序特征学习能力。为进一步强化时序表征并降低对标注数据的依赖,提出掩码分箱建模(MBM)策略用于自监督预训练,实现仅需少量标注数据即可获得鲁棒的时间表征,并作为微调阶段的辅助目标。同时,在微调中引入对比对齐与一致性正则化损失,以加强跨视角结构对应关系和时序一致性。在合成与真实世界基准上的大量实验表明,Event3R能实现鲁棒、时序一致且全局对齐的3D重建,显著优于现有事件基方法。
原文摘要 · Abstract (English)
Robust 3D reconstruction is essential for robotics and embodied perception. Recent feed-forward approaches such as DUSt3R have demonstrated impressive progress in dense 3D reconstruction from RGB images, achieving global geometric consistency and strong generalization. However, extending such dense 3D reconstruction to event cameras remains challenging due to their asynchronous, sparse, and highly dynamic nature, as well as the lack of large-scale, well-labeled datasets. In this work, we introduce Event3R, a feed-forward framework that directly maps asynchronous event streams to globally consistent 3D point clouds. Event3R represents incoming events as spatial-temporal voxels, enabling time-aware feature integration through a temporal attention module that enhances the module's temporal feature learning. To further strengthen temporal representation learning and reduce reliance on labeled data, we propose a Masked Bin Modeling (MBM) strategy for self-supervised pre-training, enabling robust temporal representation learning with minimal labeled data, and retain it as an auxiliary fine-tuning objective. In addition, contrastive alignment and consistency regularization losses are incorporated during fine-tuning to reinforce structural correspondence and temporal coherence across views. Extensive experiments on both synthetic and real-world benchmarks demonstrate that Event3R achieves robust, temporally consistent, and globally aligned 3D reconstructions, significantly outperforming existing event-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。