arXiv:2505.19487cs.CV2025-05中稿 · ICLR被引 2

用脉冲流直接估立体深,比传统方法更适应快速变化场景。

SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams

  • 基于脉冲神经网络,从原始脉冲流迭代优化深度估计。
  • 在合成与真实数据集上均超越现有方法,尤其擅长纹理缺失区域。
  • 数据效率高,少量训练数据仍保持高精度,适合低功耗设备。

传统帧式相机在快速变化场景中难以进行立体深度估计。相比之下,生物启发的脉冲相机以微秒级分辨率异步输出事件,提供一种新型感知模式。然而,现有方法缺乏针对脉冲数据设计的专用立体算法和基准。为此,我们提出SpikeStereoNet,首个直接从原始脉冲流估计立体深度的脑启发框架。该模型融合双视角原始脉冲流,并通过循环脉冲神经网络(RSNN)更新模块迭代优化深度预测。为评估方法,我们构建了大规模合成脉冲流数据集及包含稠密深度标注的真实世界立体脉冲数据集。SpikeStereoNet在两个数据集上均优于现有方法,其优势体现在对细微边缘和强度变化的捕捉能力,尤其在无纹理表面和极端光照条件下表现优异。此外,该框架具备强数据效率,在显著减少训练数据量的情况下仍保持高精度。源代码与数据集将公开发布。

原文摘要 · Abstract (English)

Conventional frame-based cameras often struggle with stereo depth estimation in rapidly changing scenes. In contrast, bio-inspired spike cameras emit asynchronous events at microsecond-level resolution, providing an alternative sensing modality. However, existing methods lack specialized stereo algorithms and benchmarks tailored to the spike data. To address this gap, we propose SpikeStereoNet, a brain-inspired framework and the first to estimate stereo depth directly from raw spike streams. The model fuses raw spike streams from two viewpoints and iteratively refines depth estimation through a recurrent spiking neural network (RSNN) update module. To benchmark our approach, we introduce a large-scale synthetic spike stream dataset and a real-world stereo spike dataset with dense depth annotations. SpikeStereoNet outperforms existing methods on both datasets by leveraging spike streams' ability to capture subtle edges and intensity shifts in challenging regions such as textureless surfaces and extreme lighting conditions. Furthermore, our framework exhibits strong data efficiency, maintaining high accuracy even with substantially reduced training data. The source code and datasets will be publicly available.

脉冲神经网络立体深度低功耗感知事件相机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。