将大模型推理能力蒸馏到轻量视觉驾驶模型,实现高效高精度自动驾驶。
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models

- 通过特征蒸馏与真实轨迹监督,压缩大模型推理能力。
- 在Bench2Drive上达到80.6分新高,超越原大模型。
- 适合追求低延迟、高能效的自动驾驶系统研发者。
利用大语言模型(LLM)的通用世界知识,有望显著提升自动驾驶系统应对罕见复杂场景的能力。尽管将LLM集成到视觉-语言-动作(VLA)模型中已取得顶尖性能,但其庞大的参数量给对延迟和能耗敏感的部署带来严峻挑战。将LLM知识蒸馏到紧凑的驾驶模型中,可在保持可管理计算开销的同时保留推理能力。此前工作多集中于简单场景和开环评估,本研究则在更复杂的交互式场景下进行闭环评估。结果表明,通过隐空间特征蒸馏与真实轨迹监督相结合,轻量级视觉仅模型Orion-Lite甚至超越其庞大的VLA教师模型ORION。在严格的Bench2Drive基准上取得80.6分的驾驶得分,创下新纪录。这揭示了纯视觉架构在高性能反应式规划中仍具有巨大未开发潜力。
原文摘要 · Abstract (English)
Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems to handle rare and complex scenarios. While integrating LLMs into Vision-Language-Action (VLA) models has yielded state-of-the-art performance, their massive parameter counts pose severe challenges for latency-sensitive and energy-efficient deployment. Distilling LLM knowledge into a compact driving model offers a compelling solution to retain these reasoning capabilities while maintaining a manageable computational footprint. Although previous works have demonstrated the efficacy of distillation, these efforts have primarily focused on relatively simple scenarios and open-loop evaluations. Therefore, in this work, we investigate LLM distillation in more complex, interactive scenarios under closed-loop evaluation. We demonstrate that through a combination of latent feature distillation and ground-truth trajectory supervision, an efficient vision-only student model \textbf{Orion-Lite} can even surpass the performance of its massive VLA teacher, ORION. Setting a new state-of-the-art on the rigorous Bench2Drive benchmark, with a Driving Score of 80.6. Ultimately, this reveals that vision-only architectures still possess significant, untapped potential for high-performance reactive planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。