arXiv:2603.10441cs.RO2026-03被引 2

用大模型理解场景,扩散模型生成可行轨迹,提升自动驾驶规划能力

KnowDiffuser: A Knowledge-Guided Diffusion Planner with LLM Reasoning

  • 大模型解析场景生成语义动作,指导扩散模型生成轨迹
  • 两阶段去噪机制使轨迹既合理又符合物理规律
  • 在nuPlan上表现超越现有方法,适合需要可解释性的规划任务

近期语言模型在语义推理方面表现出色,可用于自动驾驶的高层决策。然而,语言模型基于离散标记空间,无法生成连续且物理可行的运动轨迹。与此同时,扩散模型虽能生成可靠且动态一致的轨迹,但缺乏语义可解释性与场景理解对齐。为此,我们提出知觉引导的运动规划框架KnowDiffuser,将语言模型的语义理解与扩散模型的生成能力紧密结合。该框架利用语言模型从结构化场景表示中推断上下文感知的元动作,并将其映射为锚定后续去噪过程的先验轨迹。采用两阶段截断去噪机制高效优化轨迹,同时保持语义一致性与物理可行性。在nuPlan基准上的实验表明,KnowDiffuser在开环与闭环评估中均显著优于现有规划器,建立了鲁棒且可解释的自动驾驶系统语义-物理桥梁。

原文摘要 · Abstract (English)

Recent advancements in Language Models (LMs) have demonstrated strong semantic reasoning capabilities, enabling their application in high-level decision-making for autonomous driving (AD). However, LMs operate over discrete token spaces and lack the ability to generate continuous, physically feasible trajectories required for motion planning. Meanwhile, diffusion models have proven effective at generating reliable and dynamically consistent trajectories, but often lack semantic interpretability and alignment with scene-level understanding. To address these limitations, we propose \textbf{KnowDiffuser}, a knowledge-guided motion planning framework that tightly integrates the semantic understanding of language models with the generative power of diffusion models. The framework employs a language model to infer context-aware meta-actions from structured scene representations, which are then mapped to prior trajectories that anchor the subsequent denoising process. A two-stage truncated denoising mechanism refines these trajectories efficiently, preserving both semantic alignment and physical feasibility. Experiments on the nuPlan benchmark demonstrate that KnowDiffuser significantly outperforms existing planners in both open-loop and closed-loop evaluations, establishing a robust and interpretable framework that effectively bridges the semantic-to-physical gap in AD systems.

自动驾驶扩散模型语言模型轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。