用大模型思维链提升自动驾驶轨迹预测效率与精度
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
- 通过思维链提示让大模型生成语义标注,指导轻量模型预测
- 在5个真实数据集上超越现有方法,实现实时运行且保持强泛化能力
- 首次实现大模型能力向边缘设备轻量模型迁移,适合智能驾驶研发者
准确的运动预测对自动驾驶安全至关重要。本文提出CoT-Drive,利用大语言模型(LLMs)和思维链(CoT)提示方法增强预测性能。通过教师-学生知识蒸馏策略,将大模型的场景理解能力高效迁移到轻量语言模型,使CoT-Drive能在边缘设备实时运行,同时保持全面的场景理解与泛化能力。借助无需额外训练的CoT提示技术,模型生成的语义标注显著提升了对复杂交通环境的理解,从而增强预测的准确性和鲁棒性。此外,我们构建了两个新场景描述数据集Highway-Text和Urban-Text,用于微调轻量模型以生成上下文相关的语义注释。在五个真实世界数据集上的全面评估表明,CoT-Drive优于现有模型,展现出在复杂交通场景下的高效性与有效性。本研究首次考虑大模型在该领域的实际应用,开创性地训练并使用轻量级语言模型代理进行运动预测,树立新基准,展示大模型融入自动驾驶系统的潜力。
原文摘要 · Abstract (English)
Accurate motion forecasting is crucial for safe autonomous driving (AD). This study proposes CoT-Drive, a novel approach that enhances motion forecasting by leveraging large language models (LLMs) and a chain-of-thought (CoT) prompting method. We introduce a teacher-student knowledge distillation strategy to effectively transfer LLMs' advanced scene understanding capabilities to lightweight language models (LMs), ensuring that CoT-Drive operates in real-time on edge devices while maintaining comprehensive scene understanding and generalization capabilities. By leveraging CoT prompting techniques for LLMs without additional training, CoT-Drive generates semantic annotations that significantly improve the understanding of complex traffic environments, thereby boosting the accuracy and robustness of predictions. Additionally, we present two new scene description datasets, Highway-Text and Urban-Text, designed for fine-tuning lightweight LMs to generate context-specific semantic annotations. Comprehensive evaluations of five real-world datasets demonstrate that CoT-Drive outperforms existing models, highlighting its effectiveness and efficiency in handling complex traffic scenarios. Overall, this study is the first to consider the practical application of LLMs in this field. It pioneers the training and use of a lightweight LLM surrogate for motion forecasting, setting a new benchmark and showcasing the potential of integrating LLMs into AD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。