让AI先评判再驾驶,提升自动驾驶决策质量
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving

- 先生成粗略轨迹,再用视觉语言模型做多模态评判优化
- 在Bench2Drive上达73.33%成功率,复杂场景提升30%
- 适用于追求高安全性的自动驾驶系统研发
近期视觉语言动作(VLA)模型在自动驾驶中展现出直接将多模态输入映射为控制信号的巨大潜力。然而,以往基于VLA的方法未显式利用其作为‘评判者’的能力来优化驾驶决策,尽管该能力在其他大模型领域已得到验证,导致其在复杂闭环场景中表现受限。本文提出理论驱动的两阶段框架CriticVLA,将VLA角色从执行扩展为评判。CriticVLA首先生成初步轨迹,再通过基于VLA的评判进行多模态评估与单步优化,获得更优驾驶行为。为此,我们构建了包含1290万条标注轨迹的大规模合成数据集,覆盖多样驾驶场景,增强评判者的推理与优化能力。在Bench2Drive基准上的闭环实验表明,CriticVLA显著超越现有基线,总成功率达到73.33%,在挑战性场景中性能提升约30%。
原文摘要 · Abstract (English)
Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to control signals. However, previous VLA-based methods have not explicitly exploited the critic capability of VLAs to refine driving decisions, even though such capability has been well demonstrated in other LLM-based domains, thereby limiting their performance in complex closed-loop scenarios. In this work, we present a theoretically inspired two-stage framework, CriticVLA, which extends the role of VLAs from acting to judging. CriticVLA first generates a rough trajectory and then refines it through multimodal evaluation and single-step optimization guided by a VLA-based critic, yielding higher-quality driving behaviors. To support this process, we construct a large-scale synthetic dataset of 12.9 million annotated trajectories covering diverse driving scenarios, which enhances the critic's reasoning and refinement abilities. Extensive closed-loop experiments on the Bench2Drive benchmark show that CriticVLA significantly surpasses state-of-the-art baselines, achieving a 73.33% total success rate and delivering about 30% improvement in challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。