让动态地图支持自然语言提问和常识推理,提升人机交互效率。
Talk2DM: Enabling Natural Language Querying and Commonsense Reasoning for Vehicle-Road-Cloud Integrated Dynamic Maps with Large Language Models
- 用链式提示机制融合规则与大模型常识知识,实现自然语言查询。
- 在混合交通场景下,问答准确率超93%,响应时间仅2-5秒。
- 适配多类大模型,适合需要智能交互的自动驾驶系统开发。
动态地图(DM)是中日两国车-路-云(VRC)协同自动驾驶的基础信息基础设施。通过提供全面的交通场景表征,DM克服了独立自动驾驶系统(ADS)受物理遮挡等限制的问题。尽管基于DM的ADS已在日本实际部署,但现有系统仍缺乏支持自然语言的(NLS)人机接口,严重制约了人与地图的交互效率。为此,本文提出VRCsim——一个用于生成流式VRC协同感知(CP)数据的仿真框架,并基于此构建了聚焦空间查询与推理的VRC-QA问答数据集。在此基础上,我们进一步设计了Talk2DM,一个可即插即用的模块,为VRC-DM系统赋予自然语言查询与常识推理能力。Talk2DM采用新型链式提示(CoP)机制,逐步融合人类定义规则与大语言模型(LLMs)的常识知识。在VRC-QA上的实验表明,Talk2DM能无缝切换不同大模型并保持高精度,展现出强泛化能力。虽然大模型准确率更高,但效率显著下降。结果表明,使用Qwen3:8B、Gemma3:27B和GPT-oss模型时,Talk2DM在平均2-5秒内实现超过93%的自然语言查询准确率,具备显著实用潜力。
原文摘要 · Abstract (English)
Dynamic maps (DM) serve as the fundamental information infrastructure for vehicle-road-cloud (VRC) cooperative autonomous driving in China and Japan. By providing comprehensive traffic scene representations, DM overcome the limitations of standalone autonomous driving systems (ADS), such as physical occlusions. Although DM-enhanced ADS have been successfully deployed in real-world applications in Japan, existing DM systems still lack a natural-language-supported (NLS) human interface, which could substantially enhance human-DM interaction. To address this gap, this paper introduces VRCsim, a VRC cooperative perception (CP) simulation framework designed to generate streaming VRC-CP data. Based on VRCsim, we construct a question-answering data set, VRC-QA, focused on spatial querying and reasoning in mixed-traffic scenes. Building upon VRCsim and VRC-QA, we further propose Talk2DM, a plug-and-play module that extends VRC-DM systems with NLS querying and commonsense reasoning capabilities. Talk2DM is built upon a novel chain-of-prompt (CoP) mechanism that progressively integrates human-defined rules with the commonsense knowledge of large language models (LLMs). Experiments on VRC-QA show that Talk2DM can seamlessly switch across different LLMs while maintaining high NLS query accuracy, demonstrating strong generalization capability. Although larger models tend to achieve higher accuracy, they incur significant efficiency degradation. Our results reveal that Talk2DM, powered by Qwen3:8B, Gemma3:27B, and GPT-oss models, achieves over 93\% NLS query accuracy with an average response time of only 2-5 seconds, indicating strong practical potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。