arXiv:2510.17274cs.CV2025-10

用大语言模型提升自动驾驶行为预测,无需训练即可适配复杂场景。

Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models

  • 通过提示词从多模态大模型中提取场景结构化信息,生成可学习嵌入
  • 在Waymo和nuScenes数据集上实现持续性能提升,零样本推理不需微调
  • 适合希望快速增强现有预测模型的自动驾驶系统研发者

当前自动驾驶系统依赖专用模型进行感知与运动预测,在标准条件下表现可靠,但在多样化真实场景下的泛化能力仍面临挑战。为此,我们提出即插即用的PnF方法,将多模态大语言模型(MLLMs)融入现有运动预测模型。PnF基于自然语言能更高效描述复杂场景的洞察,通过设计提示词从MLLM中提取结构化场景理解,并将其提炼为可学习嵌入以增强行为预测模型。该方法利用MLLM的零样本推理能力,在不进行任何微调的情况下显著提升运动预测性能。我们在两个前沿运动预测模型上验证了该方法,使用Waymo Open Motion Dataset和nuScenes Dataset,结果表明在两个基准测试中均实现稳定性能提升。

原文摘要 · Abstract (English)

Current autonomous driving systems rely on specialized models for perceiving and predicting motion, which demonstrate reliable performance in standard conditions. However, generalizing cost-effectively to diverse real-world scenarios remains a significant challenge. To address this, we propose Plug-and-Forecast (PnF), a plug-and-play approach that augments existing motion forecasting models with multimodal large language models (MLLMs). PnF builds on the insight that natural language provides a more effective way to describe and handle complex scenarios, enabling quick adaptation to targeted behaviors. We design prompts to extract structured scene understanding from MLLMs and distill this information into learnable embeddings to augment existing behavior prediction models. Our method leverages the zero-shot reasoning capabilities of MLLMs to achieve significant improvements in motion prediction performance, while requiring no fine-tuning -- making it practical to adopt. We validate our approach on two state-of-the-art motion forecasting models using the Waymo Open Motion Dataset and the nuScenes Dataset, demonstrating consistent performance improvements across both benchmarks.

自动驾驶大语言模型行为预测零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。