轻量化模型实现端到端自动驾驶实时推理与可解释性。
RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation

- 通过多级监督蒸馏,将大模型能力压缩至小型学生模型。
- 视觉仅模式下推理速度提升44.8倍,视觉+语言模式下提升7.9倍。
- 支持事故关键时刻的离线语言解释,不影响实时控制。
视觉-语言-动作(VLA)模型通过联合建模视觉感知、语言推理、可解释性与动作预测,在端到端自动驾驶中展现出巨大潜力。然而,其庞大的视觉-语言主干网络和推理模块带来了显著的推理延迟,难以在道路环境中实时部署。本文提出RT-VLA,一种通过多级监督蒸馏,将最先进的SimLingo模型的驾驶与推理能力迁移到紧凑型学生模型的轻量级VLA模型。RT-VLA保持了基于语言的推理能力,并可通过离线分析安全关键驾驶时刻的语言内容实现事后解释,且不增加实时控制延迟。相较于SimLingo教师模型,RT-VLA在闭环驾驶和语言推理性能上保持竞争力,同时在仅视觉模式下推理时间减少44.8倍,在视觉+语言模式下减少7.9倍。结果表明,监督蒸馏是构建实时、可解释的VLA式自动驾驶模型的可行路径。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and action prediction. However, their large vision-language backbones and reasoning modules introduce substantial inference latency and thereby prevent their deployment in the unforgiving reality of the road networks. We propose RT-VLA, a lightweight, distilled VLA model that transfers the driving and reasoning capabilities of the state-of-the-art SimLingo model into a compact student through multi-level supervised distillation. RT-VLA preserves language-based reasoning and supports post-hoc explanation through offline language analysis of safety-critical driving moments without adding latency to real-time control. Compared to the SimLingo teacher, RT-VLA maintains competitive closed-loop driving and language reasoning performance while reducing inference time by 44.8X in vision-only mode and 7.9X in vision+language mode. These results suggest that supervised distillation is a practical approach for building real-time, explainable VLA-style autonomous driving models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。