将大模型部署到边缘设备,实现快速驾驶行为描述与推理
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
- 在路侧单元部署大模型,利用5G网络实时处理多模态数据
- 在OpenDV-Youtube数据集上实现接近人类水平的驾驶行为理解
- 通过提示工程融合环境、目标与运动信息,提升推理性能
具备强大推理能力的深度学习架构推动了自动驾驶技术的发展。将大语言模型(LLMs)应用于该领域,可在视觉任务中以接近人类感知的准确度描述驾驶场景与行为。与此同时,边缘计算因贴近数据源而日益重要,其本地化数据处理可降低传输延迟与带宽消耗,实现更快响应。本文提出一种面向边缘设备的驾驶行为描述与推理框架,由多个路侧单元构成,各单元部署LLM并借助5G NSR/NR网络通信。实验表明,部署于边缘设备的LLM可达到令人满意的响应速度。此外,我们设计了一种提示策略,整合环境、代理及运动等多模态信息以增强系统表现。在OpenDV-Youtube数据集上的实验验证了该方法在两项任务上均有显著性能提升。
原文摘要 · Abstract (English)
Deep learning architectures with powerful reasoning capabilities have driven significant advancements in autonomous driving technology. Large language models (LLMs) applied in this field can describe driving scenes and behaviors with a level of accuracy similar to human perception, particularly in visual tasks. Meanwhile, the rapid development of edge computing, with its advantage of proximity to data sources, has made edge devices increasingly important in autonomous driving. Edge devices process data locally, reducing transmission delays and bandwidth usage, and achieving faster response times. In this work, we propose a driving behavior narration and reasoning framework that applies LLMs to edge devices. The framework consists of multiple roadside units, with LLMs deployed on each unit. These roadside units collect road data and communicate via 5G NSR/NR networks. Our experiments show that LLMs deployed on edge devices can achieve satisfactory response speeds. Additionally, we propose a prompt strategy to enhance the narration and reasoning performance of the system. This strategy integrates multi-modal information, including environmental, agent, and motion data. Experiments conducted on the OpenDV-Youtube dataset demonstrate that our approach significantly improves performance across both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。