arXiv:2605.16414cs.CV2026-05

构建多传感器融合数据集,提升雷达与视觉联合目标检测性能

NERVE: A Neuromorphic Vision and Radar Ensemble for Multi-Sensor Fusion Research

论文配图:NERVE: A Neuromorphic Vision and Radar Ensemble for Multi-Sensor Fusion Research
图 1 · 摘自论文原文
  • 整合五种传感器同步数据,实现事件相机与雷达的时序对齐
  • 77GHz雷达+事件相机使检测mAP达47.5%,距离误差低于1.8米
  • 适合做多模态感知、自动驾驶与智能环境研究的开发者

我们提出NERVE(类脑视觉与雷达集成),一个包含257分钟同步数据的多传感器数据集,涵盖两个动态视觉传感器(DVS)、一个RGB-D相机及两个雷达单元(24GHz和77GHz)。数据在办公室环境中采集,共12天,包含约600GB未压缩的时序对齐数据,约91.4万帧图像及约960万条基于COCO格式的标注,覆盖16类相关物体。为评估多模态融合效果,我们构建了DVS+雷达子集用于人体检测与距离估计。基线实验表明,将DVS与77GHz雷达结合可持续提升检测性能,递归模型最高达到47.5% mAP,雷达距离均方根误差低于1.8米(以激光雷达为真值)。

原文摘要 · Abstract (English)

We present NERVE (Neuromorphic Vision and Radar Ensemble), a multi-sensor dataset comprising 257 minutes of synchronized recordings from five sensors: two Dynamic Vision Sensors (DVS), an RGB-D camera, and two Radar units (24GHz and 77GHz). Captured across 12 measurement days in office environments, NERVE contains around 600GB of uncompressed temporally aligned data with around 914,000 frames and around 9.6 million RGB COCO-formatted annotations covering 16 relevant object categories. To evaluate multi-modal fusion, we construct a DVS+Radar subset for human detection and distance estimation. Baseline experiments using feed-forward and recurrent detectors show that combining DVS with 77GHz Radar consistently improves detection, with recurrent models achieving up to 47.5% mAP and mean absolute Radar distance errors below 1.8m against LiDAR ground truth.

多传感器融合类脑视觉雷达感知目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。