构建轻量级道路场景理解基准,推动模型从感知到结构推理的升级。
RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- 设计聚焦道路拓扑与动态结构的推理任务,强调关系理解与一致性。
- 提出HRRP-T框架,通过时序一致性奖励提升视觉语言模型的几何感知能力。
- 适用于需要结构化理解的自动驾驶与数字地图构建场景。
中层道路语义理解是连接底层感知与高层规划的关键,对可靠自动驾驶和数字地图构建至关重要。现有基准多关注检测或分割等感知任务,忽视了推断道路拓扑与动态场景结构所需的推理能力。为此,我们提出RoadSceneBench——一个轻量但信息丰富的基准,用于评估复杂道路环境中的视觉推理能力。不同于大规模感知数据集,该基准强调关系理解与结构一致性,引导模型捕捉真实道路场景的内在逻辑。为增强推理可靠性,我们进一步提出层次化关系奖励传播与时间一致性(HRRP-T)训练框架,使视觉语言模型在推理过程中自适应地提升空间连贯性与语义对齐。实验表明,该方法在多种道路配置下达到领先性能。RoadSceneBench为研究中层道路语义、推动结构感知自主感知提供了紧凑而强大的基础。数据集已开源:https://github.com/XiyanLiu/RoadSceneBench。
原文摘要 · Abstract (English)
Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. However, existing benchmarks primarily target perception tasks such as detection or segmentation, overlooking the reasoning capabilities required to infer road topology and dynamic scene structure. To address this gap, we present RoadSceneBench, a lightweight yet information-rich benchmark designed to evaluate and advance visual reasoning in complex road environments. Unlike large-scale perception datasets, RoadSceneBench emphasizes relational understanding and structural consistency, encouraging models to capture the underlying logic of real-world road scenes. Furthermore, to enhance reasoning reliability, we propose Hierarchical Relational Reward Propagation with Temporal Consistency (HRRP-T), a training framework for Vision-Language Models (VLMs) in which reward signals adaptively promote spatial coherence and semantic alignment throughout the reasoning process. This paradigm enables models to move beyond static recognition toward geometry-aware and temporally consistent reasoning. Extensive experiments demonstrate that our method achieves state-of-the-art performance across diverse road configurations. RoadSceneBench thus provides a compact yet powerful foundation for studying mid-level road semantics and fostering structure-aware autonomous perception. Our dataset is available at https://github.com/XiyanLiu/RoadSceneBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。