用占据表示提升自动驾驶大模型,让车辆更懂动态环境。
Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models
- 将占据信息编码输入大模型,先分离动态与静态物体再建模。
- 4D占据预测性能提升6% IoU、4% mIoU,超越现有方法。
- 适合研究自动驾驶感知与决策融合的学者和工程师。
大型语言模型(LLMs)在机器人和自动驾驶领域取得显著进展。本文首次提出基于占据表示的大语言模型(Occ-LLM),开创性地将LLM与这一重要表征结合。为有效将占据作为输入并解决占据中的类别不平衡问题,我们提出运动分离变分自编码器(MS-VAE)。该方法利用先验知识,在输入变分自编码器前区分动态物体与静态场景,从而增强模型对动态轨迹的关注能力,并有效重建静态场景。经多项关键任务验证,包括4D占据预测、自车路径规划及基于占据的场景问答,结果表明Occ-LLM显著优于现有最先进方法,在4D占据预测任务中实现约6%的交并比(IoU)和4%的平均交并比(mIoU)提升。这些发现凸显了Occ-LLM重塑当前机器人与自动驾驶范式的变革潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode occupancy as input for the LLM and address the category imbalances associated with occupancy, we propose Motion Separation Variational Autoencoder (MS-VAE). This innovative approach utilizes prior knowledge to distinguish dynamic objects from static scenes before inputting them into a tailored Variational Autoencoder (VAE). This separation enhances the model's capacity to concentrate on dynamic trajectories while effectively reconstructing static scenes. The efficacy of Occ-LLM has been validated across key tasks, including 4D occupancy forecasting, self-ego planning, and occupancy-based scene question answering. Comprehensive evaluations demonstrate that Occ-LLM significantly surpasses existing state-of-the-art methodologies, achieving gains of about 6\% in Intersection over Union (IoU) and 4\% in mean Intersection over Union (mIoU) for the task of 4D occupancy forecasting. These findings highlight the transformative potential of Occ-LLM in reshaping current paradigms within robotic and autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。