用大模型检测自动驾驶边缘场景,提升系统安全性。
Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
- 结合开放词汇检测与提示工程,利用大模型进行上下文推理。
- 在真实失效案例中测试,验证大模型识别异常的能力。
- 适合关注自动驾驶安全与系统鲁棒性的研究者。
大语言模型(LLMs)的快速发展已拓展至多个领域。近期,研究者开始探索其在自动驾驶中的应用潜力,尤其是作为感知与规划模块的补充。然而,现有评估多基于合成数据或人工驾驶数据集,缺乏真实场景下的真值信息,难以准确衡量当前感知与规划算法的表现。为此,本文在真实边缘场景下评估了多种先进大模型,这些场景曾导致自动驾驶系统失效。所提架构结合开放词汇目标检测、提示工程与大模型上下文推理能力。实验对多个前沿模型进行了定性对比分析,并讨论了其作为自动驾驶异常检测器的潜在应用价值。
原文摘要 · Abstract (English)
The rapid evolution of large language models (LLMs) has pushed their boundaries to many applications in various domains. Recently, the research community has started to evaluate their potential adoption in autonomous vehicles and especially as complementary modules in the perception and planning software stacks. However, their evaluation is limited in synthetic datasets or manually driving datasets without the ground truth knowledge and more precisely, how the current perception and planning algorithms would perform in the cases under evaluation. For this reason, this work evaluates LLMs on real-world edge cases where current autonomous vehicles have been proven to fail. The proposed architecture consists of an open vocabulary object detector coupled with prompt engineering and large language model contextual reasoning. We evaluate several state-of-the-art models against real edge cases and provide qualitative comparison results along with a discussion on the findings for the potential application of LLMs as anomaly detectors in autonomous vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。