arXiv:2503.05164cs.ROcs.AI2025-03ICRA被引 11

用大模型构建自动驾驶行为智能评估框架,填补无综合评价方法的空白。

A Comprehensive LLM-powered Framework for Driving Intelligence Evaluation

  • 基于真人驾驶数据和访谈构建自然语言评价数据集,驱动大模型评估驾驶行为。
  • 在CARLA仿真环境中验证框架有效性,人类评估进一步确认其可靠性。
  • 适合自动驾驶算法优化与智能体设计者参考,推动更类人的自动驾驶发展。

自动驾驶的评估方法对算法优化至关重要。然而,由于驾驶智能的复杂性,目前尚无全面的自动驾驶智能水平评估方法。本文提出一种针对复杂交通环境下驾驶行为智能的评估框架,旨在填补这一空白。通过自然驾驶实验和事后行为评估访谈,构建了人类专业驾驶员与乘客的自然语言评价数据集。基于该数据集,开发了大模型驱动的驾驶评估框架。该框架在CARLA城市交通仿真器中的模拟实验中得到验证,并通过人工评估进一步证实其有效性。研究为评估和设计更智能、类人化的自动驾驶代理提供了重要参考。框架实现细节及数据集详细信息可在Github获取。

原文摘要 · Abstract (English)

Evaluation methods for autonomous driving are crucial for algorithm optimization. However, due to the complexity of driving intelligence, there is currently no comprehensive evaluation method for the level of autonomous driving intelligence. In this paper, we propose an evaluation framework for driving behavior intelligence in complex traffic environments, aiming to fill this gap. We constructed a natural language evaluation dataset of human professional drivers and passengers through naturalistic driving experiments and post-driving behavior evaluation interviews. Based on this dataset, we developed an LLM-powered driving evaluation framework. The effectiveness of this framework was validated through simulated experiments in the CARLA urban traffic simulator and further corroborated by human assessment. Our research provides valuable insights for evaluating and designing more intelligent, human-like autonomous driving agents. The implementation details of the framework and detailed information about the dataset can be found at Github.

自动驾驶大模型评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。