构建多感官交互感知数据集,助力自动驾驶实现类人智能决策
MIPD: A Multi-sensory Interactive Perception Dataset for Embodied Intelligent Driving
- 融合视觉、激光、雷达及声音、光照、振动等多模态信号
- 包含8500+帧同步标注数据,覆盖126段超20秒的复杂驾驶场景
- 适合研究具身智能驾驶的算法与多感官融合模型的团队使用
驾驶过程中,人类依赖多种感官获取信息并作出判断。为实现自动驾驶中的具身智能,必须整合多维度感官信息以增强环境交互能力。然而,现有多模态融合方案常忽略额外感官输入,制约了全自动驾驶的发展。本文提出多感官交互感知数据集MIPD,拓展当前自动驾驶算法框架,支持具身智能驾驶研究。除传统摄像头、激光雷达和4D雷达数据外,还融入声音、光照强度、振动强度和车辆速度等多种传感器信号,提升数据集全面性。数据集包含126个连续序列,多数超过20秒,涵盖超过8500帧精心同步标注的图像,覆盖多样道路与光照条件。经实验验证,该数据集为下一代自动驾驶框架探索提供了重要支持。
原文摘要 · Abstract (English)
During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional sensory information in order to facilitate interaction with the environment. However, the current multi-modal fusion sensing schemes often neglect these additional sensory inputs, hindering the realization of fully autonomous driving. This paper considers multi-sensory information and proposes a multi-modal interactive perception dataset named MIPD, enabling expanding the current autonomous driving algorithm framework, for supporting the research on embodied intelligent driving. In addition to the conventional camera, lidar, and 4D radar data, our dataset incorporates multiple sensor inputs including sound, light intensity, vibration intensity and vehicle speed to enrich the dataset comprehensiveness. Comprising 126 consecutive sequences, many exceeding twenty seconds, MIPD features over 8,500 meticulously synchronized and annotated frames. Moreover, it encompasses many challenging scenarios, covering various road and lighting conditions. The dataset has undergone thorough experimental validation, producing valuable insights for the exploration of next-generation autonomous driving frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。