arXiv:2509.10570cs.ROcs.AI2025-09综述被引 6

用大模型提升自动驾驶轨迹预测的泛化与可解释性

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

  • 用语言与场景语义融合实现可解释推理
  • 在复杂场景下显著提升预测安全性和泛化能力
  • 适合关注自动驾驶安全与大模型应用的研究者

轨迹预测是自动驾驶中的关键功能,用于预判车辆、行人等交通参与者未来的运动路径,对行车安全至关重要。传统深度学习方法虽提升了精度,但仍受限于可解释性差、依赖大规模标注数据、长尾场景泛化能力弱等问题。大型基础模型(LFMs)正改变这一研究范式,本文系统综述了近年来在轨迹预测中应用的大型语言模型(LLMs)和多模态大语言模型(MLLMs)的进展。通过融合语言与场景语义,LFMs支持可解释的上下文推理,在复杂环境中显著增强预测的安全性与泛化能力。文章总结了三种核心方法:轨迹-语言映射、多模态融合与基于约束的推理,涵盖车辆与行人预测任务、评估指标及数据集分析。讨论了计算延迟、数据稀缺与真实场景鲁棒性等关键挑战,并展望了低延迟推理、因果感知建模与运动基础模型等未来方向。

原文摘要 · Abstract (English)

Trajectory prediction serves as a critical functionality in autonomous driving, enabling the anticipation of future motion paths for traffic participants such as vehicles and pedestrians, which is essential for driving safety. Although conventional deep learning methods have improved accuracy, they remain hindered by inherent limitations, including lack of interpretability, heavy reliance on large-scale annotated data, and weak generalization in long-tail scenarios. The rise of Large Foundation Models (LFMs) is transforming the research paradigm of trajectory prediction. This survey offers a systematic review of recent advances in LFMs, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) for trajectory prediction. By integrating linguistic and scene semantics, LFMs facilitate interpretable contextual reasoning, significantly enhancing prediction safety and generalization in complex environments. The article highlights three core methodologies: trajectory-language mapping, multimodal fusion, and constraint-based reasoning. It covers prediction tasks for both vehicles and pedestrians, evaluation metrics, and dataset analyses. Key challenges such as computational latency, data scarcity, and real-world robustness are discussed, along with future research directions including low-latency inference, causality-aware modeling, and motion foundation models.

自动驾驶大模型轨迹预测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。