arXiv:2505.05360cs.RO2025-05被引 22

用轻量大模型统一自动驾驶的推理与规划,提升效率与可解释性。

DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning

  • 通过知识蒸馏将大视觉语言模型能力迁移到小型语言模型。
  • 在闭环仿真中性能媲美基准模型,且参数更小、推理更快。
  • 适合追求高效可靠端到端自动驾驶系统的研发团队使用。

我们提出DSDrive,一种面向自动驾驶的轻量化端到端框架,将推理与规划统一整合。该框架采用紧凑型大语言模型(LLM),通过知识蒸馏保留大型视觉语言模型(VLM)的增强推理能力。为有效对齐推理与规划任务,设计了基于航点驱动的双头协同模块,同步数据集结构、优化目标与学习过程。通过将两者融合于统一框架,DSDrive以规划结果为核心,融入详细推理信息,显著提升端到端系统的可解释性与可靠性。在闭环仿真中,DSDrive性能达到基准模型水平,部分关键指标更优,同时模型尺寸更小。其推理时延与内存开销显著降低,充分展现轻量化系统在可解释、高效自动驾驶中的潜力。

原文摘要 · Abstract (English)

We present DSDrive, a streamlined end-to-end paradigm tailored for integrating the reasoning and planning of autonomous vehicles into a unified framework. DSDrive leverages a compact LLM that employs a distillation method to preserve the enhanced reasoning capabilities of a larger-sized vision language model (VLM). To effectively align the reasoning and planning tasks, a waypoint-driven dual-head coordination module is further developed, which synchronizes dataset structures, optimization objectives, and the learning process. By integrating these tasks into a unified framework, DSDrive anchors on the planning results while incorporating detailed reasoning insights, thereby enhancing the interpretability and reliability of the end-to-end pipeline. DSDrive has been thoroughly tested in closed-loop simulations, where it performs on par with benchmark models and even outperforms in many key metrics, all while being more compact in size. Additionally, the computational efficiency of DSDrive (as reflected in its time and memory requirements during inference) has been significantly enhanced. Evidently thus, this work brings promising aspects and underscores the potential of lightweight systems in delivering interpretable and efficient solutions for AD.

自动驾驶大模型蒸馏端到端推理规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。