arXiv:2606.08525cs.CV2026-06被引 2

构建驾驶奖励数据集与视觉语言模型,提升自动驾驶决策智能。

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving

论文配图:DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving
图 1 · 摘自论文原文
  • 设计带时序视觉引导的驾驶轨迹评估数据集,含反事实错误行为。
  • 自研10亿参数奖励模型在特定任务上超越更大开源模型。
  • 适用于自动驾驶强化学习与多模态轨迹评分,适合算法优化者。

奖励模型在强化学习和多模态轨迹选择中至关重要。现有方法依赖人工规则或感知真值,难以泛化。尽管视觉语言模型(VLM)在其他领域展现潜力,其在驾驶任务中的表现仍不明确。本文提出DriveReward:一个通过时序视觉引导严格标注的驾驶轨迹评估数据集,并引入反事实驾驶行为增强。针对传统数据集中失败案例稀缺的问题,设计反事实标注方案以涵盖多样驾驶风格与错误行为。在该基准上的评估显示,即使领先的开源与专有VLM也未能在所有任务上表现优异,表明现有模型仍有显著改进空间。基于此,我们定制了一个10亿参数的专用奖励模型,在任务特定奖励对齐上优于更大的VLM。进一步将该模型集成至强化学习微调与多模态轨迹评分中,实现在开环与闭环评估中性能接近基于规则的奖励计算。

原文摘要 · Abstract (English)

Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such rewards typically relies on hand-crafted rule-based objectives or perception ground truth, which hinders generalization for data-scaling. While Vision-Language Models (VLMs) have demonstrated feasibility as reward models in other domains, their effectiveness in driving tasks remains underexplored. In this work, we bridge this gap by (1) introducing DriveReward, a reasoning trajectory evaluation dataset rigorously labeled via temporally-grounded visual guidance, and augmented with counterfactual driving behaviors., (2) alongside a specialized Vision-Language Reward Model. To address the scarcity of failure cases in conventional datasets, we propose a counterfactual data annotation scheme to construct cases encompassing diverse driving styles and erroneous behaviors. Evaluations on our proposed benchmark reveal that even leading open-source and proprietary VLMs fail to excel across all tasks, highlighting significant room for improvement in existing models. Building on these findings, we subsequently tailor a specialized 1B reward model that outperforms larger VLMs on task-specific reward alignment. Finally, we validate our reward model's effectiveness by integrating it into RL finetuning and multi-modal trajectory scoring across multiple baselines, achieving performance comparable to rule-based reward calculations in both open-loop and closed-loop evaluation.

自动驾驶视觉语言模型奖励模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。