arXiv:2409.03272cs.CVcs.RO2024-09被引 93

用语义占据表示构建驾驶世界模型,统一视觉语言与动作生成。

OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

  • 以语义占据为通用视觉表征,通过自回归模型融合多模态信息。
  • 在4D占据预测、路径规划和视觉问答任务上表现优异,超越现有方法。
  • 适合做自动驾驶基础模型,尤其关注场景理解与决策协同的团队。

多模态大语言模型(MLLM)的兴起推动了其在自动驾驶中的应用。现有基于MLLM的方法直接从感知映射到动作,忽视了世界动态及动作与环境变化的关系。人类则具备世界模型,能基于三维内部视觉表征模拟未来状态并规划行动。为此,我们提出OccLLaMA,一种基于语义占据的生成式世界模型,将视觉、语言、动作三者统一于自回归框架中。具体地,设计一种类VQVAE的场景分词器,高效离散化并重建稀疏且类别不平衡的语义占据场景;构建统一的多模态词汇表涵盖视觉、语言与动作;增强LLaMA模型,在统一词汇表上进行下一令牌/场景预测,实现多任务协同。大量实验表明,OccLLaMA在4D占据预测、运动规划和视觉问答任务上均达到先进水平,展现出作为自动驾驶基础模型的巨大潜力。

原文摘要 · Abstract (English)

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world and the relations between action and world dynamics. In contrast, human beings possess world model that enables them to simulate the future states based on 3D internal visual representation and plan actions accordingly. To this end, we propose OccLLaMA, an occupancy-language-action generative world model, which uses semantic occupancy as a general visual representation and unifies vision-language-action(VLA) modalities through an autoregressive model. Specifically, we introduce a novel VQVAE-like scene tokenizer to efficiently discretize and reconstruct semantic occupancy scenes, considering its sparsity and classes imbalance. Then, we build a unified multi-modal vocabulary for vision, language and action. Furthermore, we enhance LLM, specifically LLaMA, to perform the next token/scene prediction on the unified vocabulary to complete multiple tasks in autonomous driving. Extensive experiments demonstrate that OccLLaMA achieves competitive performance across multiple tasks, including 4D occupancy forecasting, motion planning, and visual question answering, showcasing its potential as a foundation model in autonomous driving.

世界模型自动驾驶多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。