arXiv:2603.09482cs.RO2026-03被引 5

让自动驾驶更懂不同驾驶风格,生成更真实可行的行驶轨迹。

StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving

  • 用物理约束+连续回归头提升轨迹可行性
  • 在1.2千场景上训练,支持5种驾驶风格
  • 轻量模型超越闭源大模型,适合真实场景落地

视觉语言模型(VLM)将视觉感知与语言推理结合。在自动驾驶中,这种融合催生了视觉语言动作(VLA)模型,可将高层多模态理解转化为驾驶行为,通常表现为未来轨迹。然而,现有VLA模型主要生成通用无碰撞轨迹。除了避障,适应多样驾驶风格(如激进、舒适)对个性化驾驶至关重要。此外,许多方法将轨迹生成视为简单标记预测,易产生运动学不可行的动作。为此,我们提出StyleVLA,一种物理感知的VLA框架,用于生成多样化且物理合理的驾驶行为。引入混合损失,结合运动学一致性约束与连续回归头,提升轨迹可行性。基于Qwen3-VL-4B构建,我们构建大规模指令数据集,包含超1.2千场景、76k鸟瞰图(BEV)样本和42k第一人称视图(FPV)样本,提供五种驾驶风格的真实轨迹及自然语言指令。实验表明,40亿参数的StyleVLA显著优于专有模型(如Gemini-3-Pro)和主流VLA模型。使用综合评分衡量成功率、物理可行性与风格契合度,StyleVLA在BEV上达0.55,在FPV上达0.51,而Gemini-3-Pro分别为0.32和0.35。结果表明,领域专用、物理感知的轻量模型可在特定任务上超越闭源大模型。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) bridge visual perception and linguistic reasoning. In Autonomous Driving (AD), this synergy has enabled Vision Language Action (VLA) models, which translate high-level multimodal understanding into driving behaviors, typically represented as future trajectories. However, existing VLA models mainly generate generic collision-free trajectories. Beyond collision avoidance, adapting to diverse driving styles (e.g., sporty, comfortable) is essential for personalized driving. Moreover, many methods treat trajectory generation as naive token prediction, which can produce kinematically infeasible actions. To address these limitations, we present StyleVLA, a physics-informed VLA framework for generating diverse and physically plausible driving behaviors. We introduce a hybrid loss that combines a kinematic consistency constraint with a continuous regression head to improve trajectory feasibility. To train StyleVLA, built on Qwen3-VL-4B, we construct a large-scale instruction dataset with over 1.2k scenarios, 76k Bird's Eye View (BEV) samples, and 42k First Person View (FPV) samples, with ground-truth trajectories for five driving styles and natural-language instructions. Experiments show that our 4B-parameter StyleVLA significantly outperforms proprietary models (e.g., Gemini-3-Pro) and state-of-the-art VLA models. Using a composite driving score measuring success rate, physical feasibility, and style adherence, StyleVLA achieves 0.55 on BEV and 0.51 on FPV, versus 0.32 and 0.35 for Gemini-3-Pro. These results show that a specialized, physics-informed, lightweight model can surpass closed-source models on domain-specific tasks.

自动驾驶视觉语言模型轨迹生成驾驶风格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。