arXiv:2601.05611cs.CV2026-01被引 4

无需语言标注,通过预测未来特征实现自动驾驶的高效决策。

FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving

  • 用自监督未来特征预测替代语言标注,让视觉语言模型直接学习驾驶表征。
  • 在NAVSIM上达到最新最佳性能,证明了预测式自监督的有效性。
  • 适合关注端到端自动驾驶与大模型应用的研究者和工程师。

尽管视觉-语言模型(VLM)为端到端自动驾驶提供了丰富的世界知识,但现有方法严重依赖人工标注的语言信息(如视觉问答)来连接感知与控制。这一范式存在离散语言标记与连续驾驶轨迹之间的根本性不匹配,常导致控制策略次优且预训练知识利用效率低下。为此,我们提出FLARE(未来感知的潜在表征),一种无需语言监督的新框架,激活预训练VLM的视觉语义能力。不同于对齐文本,我们引入自监督的未来特征预测目标,迫使模型在潜在空间中直接预测场景动态与自身运动,从而从大规模无标签轨迹数据中学习鲁棒的驾驶表示。此外,我们在规划过程中集成组相对策略优化(GRPO)以提升决策质量。在NAVSIM基准上的大量实验表明,FLARE实现了当前最优性能,验证了通过预测式自监督而非显式语言生成来利用VLM知识的有效性。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) offer rich world knowledge for end-to-end autonomous driving, current approaches heavily rely on labor-intensive language annotations (e.g., VQA) to bridge perception and control. This paradigm suffers from a fundamental mismatch between discrete linguistic tokens and continuous driving trajectories, often leading to suboptimal control policies and inefficient utilization of pre-trained knowledge. To address these challenges, we propose FLARE (Future-aware LAtent REpresentation), a novel framework that activates the visual-semantic capabilities of pre-trained VLMs without requiring language supervision. Instead of aligning with text, we introduce a self-supervised future feature prediction objective. This mechanism compels the model to anticipate scene dynamics and ego-motion directly in the latent space, enabling the learning of robust driving representations from large-scale unlabeled trajectory data. Furthermore, we integrate Group Relative Policy Optimization (GRPO) into the planning process to refine decision-making quality. Extensive experiments on the NAVSIM benchmark demonstrate that FLARE achieves state-of-the-art performance, validating the effectiveness of leveraging VLM knowledge via predictive self-supervision rather than explicit language generation.

自动驾驶视觉语言模型自监督学习端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。