arXiv:2508.02028cs.CV2025-08被引 10

构建闭环评估框架,真实测试自动驾驶视觉语言模型表现

Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving

  • 设计双系统架构,实现高阶指令到中阶动作的语义转化
  • 支持仿真与实车闭环测试,首次实现物理车辆上的交互式评估
  • 自反思场景生成可自动发现安全缺陷,适合模型安全验证

视觉语言模型(VLMs)在自动驾驶领域展现出巨大潜力,但现有评估多局限于静态输入的开环设置,忽略了能反映交互行为、反馈鲁棒性和真实安全性的闭环场景。为此,我们提出Bench2ADVLM,一个统一的分层闭环评估框架,支持在仿真与实体平台上的实时交互评估。受双过程认知理论启发,通过双系统适配架构将多种ADVLM接入仿真环境:目标ADVLM产生的异构高层驾驶指令(快速系统),由通用VLM解释为标准中阶控制动作(慢速系统),用于仿真执行。为弥合仿真与现实差距,设计物理控制抽象层,将中阶动作转化为底层执行信号,首次实现物理车辆上ADVLM的闭环测试。为进一步提升评估全面性,引入自反思场景生成模块,自动探索模型行为并发现潜在失效模式,生成高危场景。整体框架贯通高层抽象推理、中阶仿真动作与低阶现实执行。在多个前沿ADVLM与物理平台上的实验验证了该框架的诊断能力,揭示现有ADVLM在闭环条件下仍存在性能局限。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have recently emerged as a promising paradigm in autonomous driving (AD). However, current performance evaluation protocols for VLM-based AD systems (ADVLMs) are predominantly confined to open-loop settings with static inputs, neglecting the more realistic and informative closed-loop setting that captures interactive behavior, feedback resilience, and real-world safety. To address this, we introduce Bench2ADVLM, a unified hierarchical closed-loop evaluation framework for real-time, interactive assessment of ADVLMs across both simulation and physical platforms. Inspired by dual-process theories of cognition, we first adapt diverse ADVLMs to simulation environments via a dual-system adaptation architecture. In this design, heterogeneous high-level driving commands generated by target ADVLMs (fast system) are interpreted by a general-purpose VLM (slow system) into standardized mid-level control actions suitable for execution in simulation. To bridge the gap between simulation and reality, we design a physical control abstraction layer that translates these mid-level actions into low-level actuation signals, enabling, for the first time, closed-loop testing of ADVLMs on physical vehicles. To enable more comprehensive evaluation, Bench2ADVLM introduces a self-reflective scenario generation module that automatically explores model behavior and uncovers potential failure modes for safety-critical scenario generation. Overall, Bench2ADVLM establishes a hierarchical evaluation pipeline that seamlessly integrates high-level abstract reasoning, mid-level simulation actions, and low-level real-world execution. Experiments on diverse scenarios across multiple state-of-the-art ADVLMs and physical platforms validate the diagnostic strength of our framework, revealing that existing ADVLMs still exhibit limited performance under closed-loop conditions.

自动驾驶视觉语言模型闭环评估安全验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。