arXiv:2508.01057cs.AIcs.RO2025-08被引 5

用轻量视觉语言模型融合车路信息,实时规划避障路径。

Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

  • 用轻量VLM融合道路传感器与车路通信数据,实现多模态感知。
  • 在Jetson AGX Orin上实现0.57秒推理延迟,碰撞率降低77%。
  • 适合边缘部署的自动驾驶系统,提升复杂路况响应能力。

依赖车载传感器的自动驾驶系统可能无法及时发现远距离或障碍物隐患,导致可避免的碰撞;现有基于Transformer的车路协同(V2X)方法虽能缓解感知局限,却常因缺乏有效多模态融合与推理,或在高维复杂交通环境下难以满足实时性要求。本文提出实时边缘自主协驾轨迹规划框架REACT,基于微调的轻量级视觉语言模型(VLM),整合基础设施提供的危险预警与车载传感器数据,通过视觉嵌入捕捉复杂交通动态与车辆意图,解析符号输入中的精确数值,并借助上下文推理生成安全优化的行驶轨迹。为确保在边缘设备上的鲁棒实时部署,REACT创新采用残差轨迹融合(RTF)设计与专用边缘适配策略,显著降低模型复杂度并提升推理效率。在DeepAccident基准测试中,REACT达到领先性能:碰撞率降低77%,视频全景质量(VPQ)达48.2%,在Jetson AGX Orin上推理延迟仅为0.57秒。消融实验验证了各输入、模块及边缘适配策略的贡献。结果表明,轻量VLM可在边缘平台实现实时协同规划,语言引导的上下文推理对提升交通安全性与响应速度具有潜力。

原文摘要 · Abstract (English)

Autonomous driving (AD) systems relying solely on onboard sensors may fail to detect distant or obstacle hazards, potentially causing preventable collisions; however, existing transformer-based Vehicle-to-Everything (V2X) approaches, which mitigate AD sensing limitations, either lack effective multimodal fusion and reasoning or struggle to meet real-time performance requirements under complex, high-dimensional traffic conditions. This paper proposes the Real-time Edge-based Autonomous Co-pilot Trajectory planner (REACT), a V2X-integrated trajectory optimization framework for AD based on a fine-tuned lightweight Vision-Language Model (VLM). REACT integrates infrastructure-provided hazard alerts with onboard sensor data, capturing intricate surrounding traffic dynamics and vehicle intents through visual embeddings, interpreting precise numerical data from symbolic inputs, and employing contextual reasoning to generate optimized, safety-oriented trajectories. To ensure robust real-time deployment on edge devices, REACT innovatively employs Residual Trajectory Fusion (RTF) design and specialized edge-adaptation strategies to reduce model complexity and improve inference efficiency. Evaluated on the DeepAccident benchmark, REACT achieves state-of-the-art performance, a 77% collision rate reduction, a 48.2% Video Panoptic Quality (VPQ), and a 0.57-second inference latency on the Jetson AGX Orin. Ablation studies validate the contribution of each input, module, and edge adaptation strategy. These results highlight the effectiveness of lightweight VLMs in enabling real-time cooperative planning on edge platforms and underscore the potential of language-guided contextual reasoning for improving traffic safety and responsiveness.

自动驾驶多模态融合边缘计算视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。