arXiv:2506.14100cs.ROcs.SY2025-06被引 6

构建分层实测平台,评估视觉语言模型在自动驾驶中的表现

A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous Driving

  • 设计车载低延迟中间件,无缝集成多种视觉语言模型
  • 采用感知-规划-控制解耦架构,支持混合模块运行
  • 在封闭赛道实现可配置的真实场景测试,适合安全验证

视觉语言模型(VLM)通过大规模图像-文本对预训练,在自动驾驶中展现出多模态推理潜力。然而,将这类模型从通用网络数据迁移到高安全要求的驾驶场景面临显著领域偏移问题。现有基于仿真或数据集的评估方法难以充分捕捉真实场景复杂性,且难支持可重复的闭环测试与灵活场景操控。本文提出一种分层实测平台,专用于评估集成VLM的自动驾驶系统。平台包含模块化、低延迟的车载中间件,可无缝接入多种VLM;清晰分离的感知-规划-控制架构,兼容VLM与传统模块共存;以及可在封闭赛道上配置的真实测试场景,实现受控但真实的评估。通过一个集成VLM的自动驾驶案例研究,验证了该平台在多样化条件下的实验能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated notable promise in autonomous driving by offering the potential for multimodal reasoning through pretraining on extensive image-text pairs. However, adapting these models from broad web-scale data to the safety-critical context of driving presents a significant challenge, commonly referred to as domain shift. Existing simulation-based and dataset-driven evaluation methods, although valuable, often fail to capture the full complexity of real-world scenarios and cannot easily accommodate repeatable closed-loop testing with flexible scenario manipulation. In this paper, we introduce a hierarchical real-world test platform specifically designed to evaluate VLM-integrated autonomous driving systems. Our approach includes a modular, low-latency on-vehicle middleware that allows seamless incorporation of various VLMs, a clearly separated perception-planning-control architecture that can accommodate both VLM-based and conventional modules, and a configurable suite of real-world testing scenarios on a closed track that facilitates controlled yet authentic evaluations. We demonstrate the effectiveness of the proposed platform`s testing and evaluation ability with a case study involving a VLM-enabled autonomous vehicle, highlighting how our test framework supports robust experimentation under diverse conditions.

自动驾驶视觉语言模型实测平台闭环测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。