arXiv:2607.00283cs.ROcs.AI2026-07中稿 · the 2026 IEEE/RSJ …

用视觉语言模型识别影响自动驾驶决策的关键遮挡目标。

What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models

论文配图:What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
图 1 · 摘自论文原文
  • 基于规划KL散度筛选关键遮挡目标,实现精准排序。
  • 小模型微调后性能超越大模型零样本,提升约30%。
  • 适合自动驾驶感知与规划融合研究者参考。

自动驾驶车辆需在复杂环境中安全行驶,其中可能被遮挡的规划关键目标必须被识别。现有方法对所有遮挡采取统一保守策略,导致过度防御性驾驶,或忽略对规划的影响。本文提出新框架,利用信息论指标规划KL散度(PKL)系统识别并排序遮挡目标对其自身轨迹的影响。基于该排序,使用GPT-5生成包含视觉证据与推理过程的结构化标注。在nuScenes数据集上构建聚焦高影响场景的新基准。在多种通用与领域适配的VLM上进行实验,结果表明:在PKL引导数据上微调的模型性能显著提升;小模型经微调后优于大型零样本模型,且数据选择策略相比随机采样提升约30%。本工作首次系统性地训练VLM关注规划关键遮挡,实现更语义化、高效的自动驾驶风险评估。

原文摘要 · Abstract (English)

Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform conservatism, yielding needlessly defensive driving, or they infer hidden spaces without estimating the impact on the planner. This work bridges the critical gap between perception and planning by enabling Vision-Language Models (VLMs) to identify and reason about the specific hidden agents that are most critical to the ego-vehicle's trajectory. We introduce a novel framework that uses Planning KL-divergence (PKL), an information-theoretic metric, to systematically identify and rank occluded agents based on their impact on the ego vehicle's plan. Using this planning-aware ranking, we employ an expert VLM (GPT-5) to generate rich, structured annotations that capture the visual evidence and reasoning required for this task. We apply this framework to the nuScenes dataset to create a new benchmark focused on high-impact scenarios. We conduct comprehensive experiments on a wide range of general-purpose and domain-adapted VLMs, demonstrating that fine-tuning on our PKL-guided data yields dramatic performance improvements across all models. Notably, our results show that smaller, fine-tuned models significantly outperform their much larger zero-shot counterparts, and that our PKL-guided data selection strategy improves performance by approximately 30\% over random sampling. Our work presents the first systematic approach for training VLMs to focus on planning-critical occlusions, enabling more semantically grounded and efficient risk assessment in autonomous driving.

自动驾驶视觉语言模型遮挡识别规划感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。