arXiv:2502.00262cs.CVcs.AI2025-02被引 18

用视觉语言模型提升自动驾驶对罕见危险的识别能力

INSIGHT: Enhancing Autonomous Driving Safety through Vision-Language Models on Context-Aware Hazard Detection and Edge Case Evaluation

  • 构建分层视觉语言框架,融合语义与视觉信息增强场景理解
  • 在BDD100K数据集上显著提升危险预测准确率与泛化性能
  • 适合关注自动驾驶安全与边缘场景评估的研究者

自动驾驶系统在应对不可预测的边缘案例(如恶意行人动作、危险车辆行为、突发环境变化)时面临重大挑战。当前端到端驾驶模型因传统检测与预测方法局限,难以泛化至这些稀有事件。为此,我们提出INSIGHT(语义与视觉输入融合的通用危险追踪),一种分层视觉语言模型框架,旨在增强危险检测与边缘案例评估能力。通过多模态数据融合,该方法整合语义与视觉表征,实现对驾驶场景的精准解析与潜在风险的准确预测。借助对视觉语言模型的监督微调,采用基于注意力机制和坐标回归技术优化空间危险定位。在BDD100K数据集上的实验表明,该方法在危险预测的清晰度与准确性方面显著优于现有模型,泛化性能明显提升。这一进展增强了自动驾驶系统的鲁棒性与安全性,提升了复杂现实场景下的情境感知与决策能力。

原文摘要 · Abstract (English)

Autonomous driving systems face significant challenges in handling unpredictable edge-case scenarios, such as adversarial pedestrian movements, dangerous vehicle maneuvers, and sudden environmental changes. Current end-to-end driving models struggle with generalization to these rare events due to limitations in traditional detection and prediction approaches. To address this, we propose INSIGHT (Integration of Semantic and Visual Inputs for Generalized Hazard Tracking), a hierarchical vision-language model (VLM) framework designed to enhance hazard detection and edge-case evaluation. By using multimodal data fusion, our approach integrates semantic and visual representations, enabling precise interpretation of driving scenarios and accurate forecasting of potential dangers. Through supervised fine-tuning of VLMs, we optimize spatial hazard localization using attention-based mechanisms and coordinate regression techniques. Experimental results on the BDD100K dataset demonstrate a substantial improvement in hazard prediction straightforwardness and accuracy over existing models, achieving a notable increase in generalization performance. This advancement enhances the robustness and safety of autonomous driving systems, ensuring improved situational awareness and potential decision-making in complex real-world scenarios.

自动驾驶视觉语言模型危险检测边缘案例

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。