提出快速通用的自动驾驶视觉语言动作模型,支持高效推理与跨场景泛化。
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
- 用可学习的动作查询并行生成连续驾驶轨迹,基于训练数据初始化。
- 在多个基准上达到当前最佳性能,推理速度领先,泛化能力显著。
- 整合8个公开数据集,统一为链式思维格式,适合新场景和车辆配置迁移。
视觉-语言-动作(VLA)模型在自动驾驶决策中展现出强大能力,但现有模型常面临推理效率低、难以泛化至新车辆配置与驾驶场景的问题。本文提出Reasoning-VLA,一种通用且高效的动作生成VLA框架。该模型采用一组可学习的动作查询,通过从训练语料库中的真实轨迹进行高斯采样初始化,与增强的视觉-语言特征交互,实现连续动作轨迹的并行生成。为提升泛化能力,我们整合了八个公开的自动驾驶数据集,构建标准化、基于链式思维推理的数据格式,便于模型训练。结合监督学习与强化学习微调,大量实证评估表明,Reasoning-VLA在多个基准测试中达到当前最优性能,具备卓越的泛化能力与迄今为止最快的推理速度。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Reasoning-VLA, a general and fast action-generation VLA framework. The proposed model employs a set of learnable action queries, initialized via Gaussian sampling from ground-truth trajectories within the training corpus. These learnable queries interact with reasoning-enhanced vision-language features to generate continuous action trajectories in parallel. To promote robust generalization, we consolidate eight publicly available autonomous driving datasets into a standardized, Chain-of-Thought reasoning-based, and easy-to-use data format for model training. Leveraging both supervised learning and reinforcement learning fine-tuning, extensive empirical evaluations across multiple benchmarks demonstrate that Reasoning-VLA achieves state-of-the-art performance, superior generalization capability, and the excellent inference speed reported to date.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。