arXiv:2409.15176cs.CV2024-09中稿 · ACCV 2024被引 4

用脉冲流训练3D高斯场,实现实时高质量视图合成。

SpikeGS: Learning 3D Gaussian Fields from Continuous Spike Stream

  • 基于3DGS设计可微分脉冲渲染框架,融合噪声嵌入与脉冲神经元。
  • 在极低光噪声下仍能还原精细纹理,实时渲染速度优于现有方法。
  • 适合脉冲相机视觉任务,尤其适用于高速动态场景重建。

脉冲相机是一种高速视觉传感器,相比传统帧相机具有更高的时间分辨率和动态范围,在诸多计算机视觉任务中表现优异。然而,基于脉冲相机的新型视图合成研究仍不充分。尽管已有方法尝试从脉冲流学习神经辐射场,但它们在极端噪声、低质量光照条件下缺乏鲁棒性,或因使用深度全连接网络与射线追踪渲染策略导致计算复杂度高,难以恢复精细纹理。相比之下,最新的3DGS通过将点云优化为高斯椭球,实现了高质量实时渲染。受此启发,我们提出SpikeGS,一种仅从脉冲流学习3D高斯场的方法。我们设计了一种基于3DGS的可微分脉冲流渲染框架,引入噪声嵌入与脉冲神经元。借助3DGS的多视角一致性及基于瓦片的多线程并行渲染机制,实现了高质量实时渲染。此外,我们提出了一个泛化能力强的脉冲渲染损失函数,适应不同光照条件。实验表明,该方法能从移动脉冲相机捕获的连续脉冲流中重建出具有精细纹理的视图合成结果,在极低光噪声环境下表现出高鲁棒性。真实与合成数据集上的实验均证明,本方法在渲染质量与速度上均超越现有方法。

原文摘要 · Abstract (English)

A spike camera is a specialized high-speed visual sensor that offers advantages such as high temporal resolution and high dynamic range compared to conventional frame cameras. These features provide the camera with significant advantages in many computer vision tasks. However, the tasks of novel view synthesis based on spike cameras remain underdeveloped. Although there are existing methods for learning neural radiance fields from spike stream, they either lack robustness in extremely noisy, low-quality lighting conditions or suffer from high computational complexity due to the deep fully connected neural networks and ray marching rendering strategies used in neural radiance fields, making it difficult to recover fine texture details. In contrast, the latest advancements in 3DGS have achieved high-quality real-time rendering by optimizing the point cloud representation into Gaussian ellipsoids. Building on this, we introduce SpikeGS, the method to learn 3D Gaussian fields solely from spike stream. We designed a differentiable spike stream rendering framework based on 3DGS, incorporating noise embedding and spiking neurons. By leveraging the multi-view consistency of 3DGS and the tile-based multi-threaded parallel rendering mechanism, we achieved high-quality real-time rendering results. Additionally, we introduced a spike rendering loss function that generalizes under varying illumination conditions. Our method can reconstruct view synthesis results with fine texture details from a continuous spike stream captured by a moving spike camera, while demonstrating high robustness in extremely noisy low-light scenarios. Experimental results on both real and synthetic datasets demonstrate that our method surpasses existing approaches in terms of rendering quality and speed.

3D高斯脉冲相机视图合成实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。