arXiv:2505.02593cs.CV2025-05CVPR

融合事件相机与激光雷达数据,用注意力机制实现高精度稠密深度估计

DELTA: Dense Depth from Events and LiDAR using Transformer's Attention

论文配图:DELTA: Dense Depth from Events and LiDAR using Transformer's Attention
图 1 · 摘自论文原文
  • 利用自注意力和交叉注意力建模事件与激光雷达的时空关联
  • 在近距离场景下误差降低至之前最先进方法的1/4
  • 适合自动驾驶、机器人等需要高精度实时深度感知的场景

事件相机和激光雷达提供互补但不同的数据:前者异步检测光照变化,后者以固定频率提供稀疏但精确的深度信息。目前极少研究探索这两种模态的结合。本文提出一种基于神经网络的新方法,用于融合事件与激光雷达数据以估计稠密深度图。所提出的架构DELTA利用自注意力和交叉注意力机制,建模事件与激光雷达数据内部及之间的空间与时间关系。经过全面评估,结果表明DELTA在事件驱动深度估计任务上达到新最优性能,且在近距离范围内误差相比先前最先进方法最多降低四倍。

原文摘要 · Abstract (English)

Event cameras and LiDARs provide complementary yet distinct data: respectively, asynchronous detections of changes in lighting versus sparse but accurate depth information at a fixed rate. To this day, few works have explored the combination of these two modalities. In this article, we propose a novel neural-network-based method for fusing event and LiDAR data in order to estimate dense depth maps. Our architecture, DELTA, exploits the concepts of self- and cross-attention to model the spatial and temporal relations within and between the event and LiDAR data. Following a thorough evaluation, we demonstrate that DELTA sets a new state of the art in the event-based depth estimation problem, and that it is able to reduce the errors up to four times for close ranges compared to the previous SOTA.

深度估计多模态融合注意力机制事件相机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。