arXiv:2605.20942cs.CV2026-05

用图结构统一道路几何与语义,让自动驾驶理解更精准。

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding

论文配图:Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding
图 1 · 摘自论文原文
  • 构建道路图谱统一几何与语言信息,支持复杂问答生成
  • 仅需20-80个场景训练小模型,即在组合推理上显著提升
  • 适合自动驾驶系统、视觉语言模型研究者使用

车道几何、拓扑关系及交通元素间的关系结构化理解是安全自动驾驶的基础。尽管视觉语言模型(VLMs)具备良好的语义灵活性,但缺乏精确的几何与关系支撑;而传统模块化系统(如HD地图、拓扑道路图)虽具结构精度,却语义僵化。为此,我们提出联合道路基底(CRS),一种基于图的框架,使几何结构与开放词汇语义能在单一表示中协同执行。CRS通过递归图查询自动生成复合且语言多样的问题-答案对,并引入“免费定位”机制确保逻辑可追溯至具体地图元素,同时程序化提取思维链监督信号。实验表明,当前主流VLMs在结构化道路推理中表现不佳,而仅用20至80个CRS增强场景训练20亿或40亿参数的小模型,即可在不同深度的组合推理任务中获得稳定提升。通过可验证推理轨迹分析发现,基线模型失败于关系理解,而CRS训练模型的失败主要集中在属性识别,说明道路理解的核心瓶颈并非模型规模,而是缺乏结构化监督。

原文摘要 · Abstract (English)

Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural precision but remain semantically rigid. To bridge this gap, we introduce the Combined Road Substrate (CRS), a graph-grounded framework that makes geometric road structure and open-vocabulary semantics jointly executable in a single representation. CRS enables the automatic generation of compositionally complex and linguistically varied question-answer pairs via recursive graph queries, augmented with a "grounding for free" mechanism that ensures logical traceability to specific map elements, and procedurally extracted chain-of-thought supervision traces. We demonstrate that state-of-the-art VLMs - including large, closed-source models - struggle significantly with structured road reasoning, yet training a small 2- or 4-billion-parameter model with as few as 20 to 80 CRS-enriched scenes yields stable gains in compositional reasoning tasks of varying depth. Analysis of model behavior via verifiable reasoning traces reveals a systematic shift in failure modes: whereas baseline models fail at relational scene understanding, CRS-trained models reduce failures to attribute recognition, suggesting that the primary bottleneck in road understanding is not model scale, but the absence of structured supervision.

自动驾驶图神经网络视觉语言模型结构化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。