arXiv:2604.03685cs.CV2026-04中稿 · CVPR被引 1

构建多模态驾驶数据集,提升复杂环境感知鲁棒性

DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR

  • 融合双目事件、可见光、热成像与4D雷达、双激光雷达
  • 提供2D/3D边界框与轨迹信息,支持跨传感器对比
  • 提出统一特征空间融合框架,增强恶劣天气检测能力

本文提出DSERT-RoLL,一个包含双目事件相机、可见光相机、热成像相机、4D雷达和双激光雷达的驾驶数据集,覆盖多种天气与光照条件。数据集提供精确的2D与3D边界框、轨迹ID及自车位姿,支持不同传感器组合间的公平比较。旨在缓解事件相机与4D雷达等新型传感器的数据稀缺问题,并推动其行为的系统性研究。建立统一的2D与3D基准,实现跨传感器族与族内性能的直接对比。报告了代表性单模态与多模态方法的基线结果,提供促进融合策略与传感器组合研究的评估协议。此外,提出一种将各传感器特有线索融入统一特征空间的融合框架,在多样天气与光照条件下提升了3D检测鲁棒性。

原文摘要 · Abstract (English)

In this paper, we present DSERT-RoLL, a driving dataset that incorporates stereo event, RGB, and thermal cameras together with 4D radar and dual LiDAR, collected across diverse weather and illumination conditions. The dataset provides precise 2D and 3D bounding boxes with track IDs and ego vehicle odometry, enabling fair comparisons within and across sensor combinations. It is designed to alleviate data scarcity for novel sensors such as event cameras and 4D radar and to support systematic studies of their behavior. We establish unified 3D and 2D benchmarks that enable direct comparison of characteristics and strengths across sensor families and within each family. We report baselines for representative single modality and multimodal methods and provide protocols that encourage research on different fusion strategies and sensor combinations. In addition, we propose a fusion framework that integrates sensor specific cues into a unified feature space and improves 3D detection robustness under varied weather and lighting.

多模态感知自动驾驶传感器融合鲁棒检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。