arXiv:2606.18242cs.CV2026-06

用事件相机提升自动驾驶的感知与决策能力

EventDrive: Event Cameras for Vision-Language Driving Intelligence

论文配图:EventDrive: Event Cameras for Vision-Language Driving Intelligence
图 1 · 摘自论文原文
  • 融合事件流、图像帧与语言指令,构建多任务驾驶理解框架
  • 事件数据显著提升运动感知精度与复杂场景鲁棒性
  • 适合研究自动驾驶多模态融合与实时决策的学者

事件相机通过微秒级延迟和高动态范围的异步亮度变化感知世界,其运动保真度远超传统帧式传感器,能捕捉常规曝光遗漏的时间结构。这些特性使事件数据成为自动驾驶中RGB的有力补充,尤其在模糊、强光和快速运动场景下,帧式感知易失效。然而现有事件感知视觉语言模型仍局限于通用感知,未揭示事件传感如何支持全链路驾驶推理与决策。本文提出EventDrive,一个涵盖感知、理解、预测与规划四大维度的大规模基准与模型套件,包含描述生成、结构化问答、定位、运动状态识别、轨迹预测与路径规划等任务。在此基础上,EventDrive-VLM引入多时程事件金字塔与时间-时程混合专家模块,自适应编码并融合异步事件与帧式信息以支持下游推理。全面评估表明,事件流显著提升了时间精度、运动意识与鲁棒性,将事件感知置于驾驶智能的核心位置。

原文摘要 · Abstract (English)

Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement to RGB in autonomous driving, especially under blur, glare, and rapid motion, where frame-based perception can become unreliable. However, existing event-aware vision-language models remain limited to generic perception and do not reveal how event sensing contributes to reasoning and decision-making across the full driving loop. We present EventDrive, a large-scale benchmark and model suite that unifies event streams, RGB frames, and language supervision across four core dimensions: Perception, Understanding, Prediction, and Planning, covering captions, structured QA, grounding, motion-state recognition, trajectory forecasting, and planning tasks. Building on this foundation, EventDrive-VLM introduces a multi-horizon event pyramid and a temporal-horizon mixture-of-experts module to adaptively encode and fuse asynchronous and frame-based information for downstream reasoning. Comprehensive evaluation across diverse tasks shows that event streams provide substantial gains in temporal precision, motion awareness, and robustness, bringing event sensing into the center of driving intelligence.

事件相机自动驾驶多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。