arXiv:2503.18730cs.CL2025-03被引 7

用知识图谱构建驾驶场景理解模型,提升未来场景预测能力。

Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving

  • 基于知识图谱构建鸟瞰符号表示,融合道路拓扑与交通规则。
  • 微调T5模型实现86.7%的未来场景预测准确率。
  • 适合自动驾驶场景理解、多智能体交互研究者参考。

自动驾驶领域在目标识别、轨迹预测和运动规划等方面取得了显著进展,但现有方法在有效理解驾驶场景随时间演化的复杂性方面仍存在局限。本文提出FM4SU,一种用于自动驾驶场景理解的符号基础模型(FM)训练新方法。该方法利用知识图谱(KG)捕获感官观测及领域知识,如道路拓扑、交通规则或交通参与者间的复杂交互。从每个驾驶场景的KG中提取鸟瞰视图(BEV)符号表示,包含对象间的时空信息。该表示被序列化为令牌序列,输入预训练语言模型(PLMs)以学习驾驶场景元素间的共现规律并生成未来场景预测。我们在nuScenes数据集上进行了多种场景下的实验。结果表明,微调后的模型在所有任务中均显著提升准确率,其中微调的T5模型在下一场景预测任务中达到86.7%的准确率。论文结论认为,FM4SU为构建更全面的自动驾驶场景理解模型提供了有前景的基础。

原文摘要 · Abstract (English)

The autonomous driving field has seen remarkable advancements in various topics, such as object recognition, trajectory prediction, and motion planning. However, current approaches face limitations in effectively comprehending the complex evolutions of driving scenes over time. This paper proposes FM4SU, a novel methodology for training a symbolic foundation model (FM) for scene understanding in autonomous driving. It leverages knowledge graphs (KGs) to capture sensory observation along with domain knowledge such as road topology, traffic rules, or complex interactions between traffic participants. A bird's eye view (BEV) symbolic representation is extracted from the KG for each driving scene, including the spatio-temporal information among the objects across the scenes. The BEV representation is serialized into a sequence of tokens and given to pre-trained language models (PLMs) for learning an inherent understanding of the co-occurrence among driving scene elements and generating predictions on the next scenes. We conducted a number of experiments using the nuScenes dataset and KG in various scenarios. The results demonstrate that fine-tuned models achieve significantly higher accuracy in all tasks. The fine-tuned T5 model achieved a next scene prediction accuracy of 86.7%. This paper concludes that FM4SU offers a promising foundation for developing more comprehensive models for scene understanding in autonomous driving.

自动驾驶知识图谱场景理解预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。