融合事件与图像信息,实现低功耗下更精准的车辆转向预测。
Energy-Aware Imitation Learning for Steering Prediction Using Events and Frames
- 设计能量驱动的跨模态融合模块,动态整合事件与帧数据。
- 在DDD20和DRFuser数据集上,转向预测误差低于现有SOTA方法。
- 适合对能耗敏感的自动驾驶系统,尤其适用于复杂光照场景。
在自动驾驶中,仅依赖基于帧的摄像头可能导致因长曝光、高速运动和恶劣光照条件带来的精度下降。为解决这一问题,本文引入一种类生物视觉传感器——事件相机。与传统相机不同,事件相机以稀疏、异步的方式捕捉事件,可作为互补模态缓解上述挑战。本文提出一种面向转向预测的能量感知模仿学习框架,融合事件与帧信息。具体设计了能量驱动的跨模态融合模块(ECFM)与能量感知解码器,以生成可靠且安全的预测结果。在两个公开的真实世界数据集DDD20与DRFuser上的大量实验表明,本方法优于现有最先进(SOTA)方法。代码与训练模型将在论文录用后发布。
原文摘要 · Abstract (English)
In autonomous driving, relying solely on frame-based cameras can lead to inaccuracies caused by factors like long exposure times, high-speed motion, and challenging lighting conditions. To address these issues, we introduce a bio-inspired vision sensor known as the event camera. Unlike conventional cameras, event cameras capture sparse, asynchronous events that provide a complementary modality to mitigate these challenges. In this work, we propose an energy-aware imitation learning framework for steering prediction that leverages both events and frames. Specifically, we design an Energy-driven Cross-modality Fusion Module (ECFM) and an energy-aware decoder to produce reliable and safe predictions. Extensive experiments on two public real-world datasets, DDD20 and DRFuser, demonstrate that our method outperforms existing state-of-the-art (SOTA) approaches. The codes and trained models will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。