构建真实密集交通中车辆汇入的闭环评测基准,提升评估真实性与多样性。
Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive Vehicles
- 用大规模数据训练的微观交互车辆模拟真实交通行为
- 引入大语言模型评估自动驾驶车辆汇入表现,超越传统指标
- 适用于评估复杂交互场景下的智能驾驶规划算法
尽管自动驾驶能力快速进步,但在密集交通中汇入仍是重大挑战。现有运动规划方法缺乏有效评估手段:多数封闭回路仿真依赖规则控制其他车辆,导致行为单一、随机性不足,难以真实反映高度交互场景下的规划能力。同时,传统评价指标无法全面衡量汇入性能。为此,我们提出一个闭环评测基准,用于评估车辆汇入场景中的运动规划能力。该基准通过大规模数据集训练的具备微观行为特征的其他车辆,显著提升了场景复杂度与多样性。此外,采用大语言模型(LLMs)重构评估机制,对自主车辆汇入主路的表现进行综合评判。大量实验与实车测试验证了该基准的先进性。基于此,我们对现有方法进行了评估,识别出共性问题。仿真环境与评估流程可访问 https://github.com/WZM5853/Bench4Merge。
原文摘要 · Abstract (English)
While the capabilities of autonomous driving have advanced rapidly, merging into dense traffic remains a significant challenge, many motion planning methods for this scenario have been proposed but it is hard to evaluate them. Most existing closed-loop simulators rely on rule-based controls for other vehicles, which results in a lack of diversity and randomness, thus failing to accurately assess the motion planning capabilities in highly interactive scenarios. Moreover, traditional evaluation metrics are insufficient for comprehensively evaluating the performance of merging in dense traffic. In response, we proposed a closed-loop evaluation benchmark for assessing motion planning capabilities in merging scenarios. Our approach involves other vehicles trained in large scale datasets with micro-behavioral characteristics that significantly enhance the complexity and diversity. Additionally, we have restructured the evaluation mechanism by leveraging Large Language Models (LLMs) to assess each autonomous vehicle merging onto the main lane. Extensive experiments and test-vehicle deployment have demonstrated the progressiveness of this benchmark. Through this benchmark, we have obtained an evaluation of existing methods and identified common issues. The simulation environment and evaluation process can be accessed at https://github.com/WZM5853/Bench4Merge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。