用事件相机实现带不确定性的单目深度估计,提升可靠性。
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling

- 基于事件流构建深度分布,融合高斯、对数正态和证据学习估算不确定性。
- 10帧对数正态与5帧证据学习在误差指标上表现最佳。
- 适合做事件相机感知、自动驾驶中需可信深度的场景。
事件相机相比传统帧式传感器具有微秒级时间分辨率、高动态范围和低带宽优势。本文利用深度神经网络从单目事件流中预测像素级深度分布,并采用高斯、对数正态及证据学习框架进行不确定性建模。对比了六种事件表示方法:含1、5、10、20个时间桶的时空体素网格、紧凑时空表示(CSTR)和时间有序近期事件(TORE)体素。基于U-Net的模型先在合成数据上训练,再在真实序列上微调。评估指标包括绝对相对误差、均方根误差及稀疏化误差曲线下面积。定量结果显示各表示性能相近,其中10桶对数正态与5桶证据学习在各项指标上最优。实验表明不确定性估计可成功融入事件驱动单目深度估计,并有效标识可靠深度像素。
原文摘要 · Abstract (English)
Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and low bandwidth. In this paper, we predict per-pixel depth distributions from monocular event streams using deep neural networks. We estimate uncertainty using Gaussian, log-normal, and evidential learning frameworks. We compare six event representations: spatio-temporal voxel grids with 1, 5, 10, and 20 temporal bins, the Compact Spatio-Temporal Representation (CSTR), and Time-Ordered Recent Event (TORE) volumes. Our U-Net-based models are trained on synthetic data and then fine-tuned on real sequences. We evaluate performance using absolute relative error, root mean squared error, and the area under the sparsification error. Quantitative results show that the representations perform similarly, while 10 bin log-normal and 5 bin evidential learning perform best across metrics. Our experiments demonstrate that uncertainty estimation can be successfully integrated into event-based monocular depth estimation, and be used to indicate pixels with reliable depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。