用事件相机生成3D场景,通过视频先验提升重建质量
Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video Priors
- 用预训练事件转视频模型生成初始帧,再用事件流细化细节
- 逐步减少监督事件数,缓解事件稀疏带来的优化困难
- 适合做高速、低光等极端条件下的3D重建任务
事件相机是生物启发式传感器,输出异步稀疏事件流而非固定帧。得益于高动态范围和高时间分辨率等优势,事件相机已被用于机器人地图构建中的3D重建。近期神经渲染技术如3D高斯点云(3DGS)在3D重建中表现优异,但如何构建有效的基于事件的3DGS流程仍待探索。由于3DGS通常依赖高质量初始化和密集多视角约束,而事件本身具有固有的稀疏性,导致优化困难。为此,我们提出新型事件驱动3DGS框架Elite-EvGS。核心思路是从现成的事件到视频(E2V)模型中蒸馏先验知识,实现粗到细的3D场景重建。具体而言,针对3DGS初始化复杂性,提出一种新颖的预热初始化策略:先用E2V模型生成的帧优化粗粒度3DGS,再引入事件流精修细节;同时设计渐进式事件监督策略,通过窗口切片操作逐步减少用于监督的事件数量,缓解事件帧的时间随机性,有利于局部纹理与全局结构的优化。在基准数据集上的实验表明,Elite-EvGS能重建出更精细的纹理与结构。同时,在真实世界数据上也表现出良好性能,涵盖快速运动和低光照等挑战性场景。
原文摘要 · Abstract (English)
Event cameras are bio-inspired sensors that output asynchronous and sparse event streams, instead of fixed frames. Benefiting from their distinct advantages, such as high dynamic range and high temporal resolution, event cameras have been applied to address 3D reconstruction, important for robotic mapping. Recently, neural rendering techniques, such as 3D Gaussian splatting (3DGS), have been shown successful in 3D reconstruction. However, it still remains under-explored how to develop an effective event-based 3DGS pipeline. In particular, as 3DGS typically depends on high-quality initialization and dense multiview constraints, a potential problem appears for the 3DGS optimization with events given its inherent sparse property. To this end, we propose a novel event-based 3DGS framework, named Elite-EvGS. Our key idea is to distill the prior knowledge from the off-the-shelf event-to-video (E2V) models to effectively reconstruct 3D scenes from events in a coarse-to-fine optimization manner. Specifically, to address the complexity of 3DGS initialization from events, we introduce a novel warm-up initialization strategy that optimizes a coarse 3DGS from the frames generated by E2V models and then incorporates events to refine the details. Then, we propose a progressive event supervision strategy that employs the window-slicing operation to progressively reduce the number of events used for supervision. This subtly relives the temporal randomness of the event frames, benefiting the optimization of local textural and global structural details. Experiments on the benchmark datasets demonstrate that Elite-EvGS can reconstruct 3D scenes with better textural and structural details. Meanwhile, our method yields plausible performance on the captured real-world data, including diverse challenging conditions, such as fast motion and low light scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。