arXiv:2512.09864cs.CV2025-12被引 14

统一理解生成与规划,提升自动驾驶长尾场景表现

UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving

  • 融合视觉、语言与动作,构建端到端协同框架
  • 在多个数据集上实现感知、推理与决策最优性能
  • 适合研究自动驾驶多模态融合与复杂场景泛化

自动驾驶系统在长尾场景中表现受限,源于世界知识不足和视觉动态建模能力弱。现有基于视觉-语言-动作(VLA)的方法无法利用未标注视频进行视觉因果学习,而世界模型方法缺乏大语言模型的推理能力。本文构建了多个专用数据集,包含复杂场景下的推理与规划标注。提出统一的理解-生成-规划框架UniUGP,通过混合专家架构协同场景推理、未来视频生成与轨迹规划。集成预训练视觉语言模型与视频生成模型,利用视觉动态与语义推理提升规划性能。以多帧观测与语言指令为输入,输出可解释的思维链、物理一致的轨迹与连贯的未来视频。采用四阶段训练策略,在多个现有自动驾驶数据集及自建数据集上逐步构建能力。实验表明,该方法在感知、推理与决策任务上达到当前最优水平,对挑战性长尾场景具有更强泛化能力。

原文摘要 · Abstract (English)

Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for visual causal learning, while world model-based methods lack reasoning capabilities from large language models. In this paper, we construct multiple specialized datasets providing reasoning and planning annotations for complex scenarios. Then, a unified Understanding-Generation-Planning framework, named UniUGP, is proposed to synergize scene reasoning, future video generation, and trajectory planning through a hybrid expert architecture. By integrating pre-trained VLMs and video generation models, UniUGP leverages visual dynamics and semantic reasoning to enhance planning performance. Taking multi-frame observations and language instructions as input, it produces interpretable chain-of-thought reasoning, physically consistent trajectories, and coherent future videos. We introduce a four-stage training strategy that progressively builds these capabilities across multiple existing AD datasets, along with the proposed specialized datasets. Experiments demonstrate state-of-the-art performance in perception, reasoning, and decision-making, with superior generalization to challenging long-tail situations.

自动驾驶多模态规划生成长尾场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。