arXiv:2602.07680cs.CVcs.AI2026-02被引 1

用视觉语言模型提升自动驾驶安全评估与规划,让车更懂语义风险。

Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning

  • 用CLIP做轻量级无类别危险检测,快速识别未知路障。
  • 全局语义嵌入不能直接提升轨迹精度,需任务对齐提取方法。
  • 用自然语言指令约束行为,减少极端事故风险,适合复杂场景。

视觉语言模型(VLMs)作为强大的表示学习系统,能对齐视觉感知与自然语言概念,为自动驾驶中的语义推理带来新机遇。本文研究将视觉语言表示集成到感知、预测和规划流程中,支持驾驶场景安全评估与决策。第一,提出一种基于CLIP的轻量级、无类别危害筛查方法,利用图像-文本相似性生成低延迟语义危险信号,无需显式目标检测或视觉问答即可鲁棒检测多样且分布外的道路隐患。第二,在Waymo Open Dataset上,将场景级视觉语言嵌入引入基于Transformer的轨迹规划框架,结果表明直接以全局嵌入条件化规划器无法提升轨迹准确性,凸显表示-任务对齐的重要性,推动任务感知提取方法的发展。第三,利用doScenes数据集,探索自然语言作为运动规划中的显式行为约束,乘客风格指令基于视觉场景元素可抑制罕见但严重的规划失败,在模糊场景中改善安全对齐行为。综合来看,当用于表达语义风险、意图和行为约束时,视觉语言表示在自动驾驶安全方面具有巨大潜力。实现这一潜力本质上是工程问题,需精细系统设计与结构化语义接地,而非简单特征注入。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous driving. This paper investigates how vision-language representations support driving scene safety assessment and decision-making when integrated into perception, prediction, and planning pipelines. We study three complementary system-level use cases. First, we introduce a lightweight, category-agnostic hazard screening approach leveraging CLIP-based image-text similarity to produce a low-latency semantic hazard signal. This enables robust detection of diverse and out-of-distribution road hazards without explicit object detection or visual question answering. Second, we examine the integration of scene-level vision-language embeddings into a transformer-based trajectory planning framework using the Waymo Open Dataset. Our results show that naively conditioning planners on global embeddings does not improve trajectory accuracy, highlighting the importance of representation-task alignment and motivating the development of task-informed extraction methods for safety-critical planning. Third, we investigate natural language as an explicit behavioral constraint on motion planning using the doScenes dataset. In this setting, passenger-style instructions grounded in visual scene elements suppress rare but severe planning failures and improve safety-aligned behavior in ambiguous scenarios. Taken together, these findings demonstrate that vision-language representations hold significant promise for autonomous driving safety when used to express semantic risk, intent, and behavioral constraints. Realizing this potential is fundamentally an engineering problem requiring careful system design and structured grounding rather than direct feature injection.

视觉语言自动驾驶安全评估语义表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。