arXiv:2601.16964cs.AI2026-01被引 4

构建30万条自动驾驶场景数据集,用于训练和评估大模型决策能力。

AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems

  • 用大模型生成多样化驾驶场景,覆盖7个独立变量维度。
  • 测试50个主流大模型,发现开源模型在物理推理上快速追赶闭源模型。
  • 适合研究自主系统智能、大模型泛化能力的学者与开发者。

大语言模型(LLMs)在自主系统中用于推理驱动的感知、规划与决策,但缺乏大规模、结构化且安全关键的评测基准。本文提出AgentDrive,一个包含30万条由大模型生成的驾驶场景的开放基准数据集,用于训练、微调和评估自主代理在多样条件下的表现。该数据集在七个正交维度上定义了场景空间:场景类型、驾驶员行为、环境、道路布局、目标、难度与交通密度。通过大模型驱动的提示转JSON管道生成语义丰富、可仿真且符合物理与模式约束的场景规范。每个场景经过仿真推演、替代安全度量计算及规则化结果标注。为补充仿真评估,我们引入AgentDrive-MCQ——一个涵盖五个推理维度(物理、策略、混合、场景、比较)的10万道多选题基准。对50个领先大模型的大规模评估显示,尽管专有前沿模型在上下文与策略推理上最优,但先进开源模型在结构化与物理基础推理方面正迅速缩小差距。相关数据集、基准、评估代码与材料已开源至https://github.com/maferrag/AgentDrive。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has sparked growing interest in their integration into autonomous systems for reasoning-driven perception, planning, and decision-making. However, evaluating and training such agentic AI models remains challenging due to the lack of large-scale, structured, and safety-critical benchmarks. This paper introduces AgentDrive, an open benchmark dataset containing 300,000 LLM-generated driving scenarios designed for training, fine-tuning, and evaluating autonomous agents under diverse conditions. AgentDrive formalizes a factorized scenario space across seven orthogonal axes: scenario type, driver behavior, environment, road layout, objective, difficulty, and traffic density. An LLM-driven prompt-to-JSON pipeline generates semantically rich, simulation-ready specifications that are validated against physical and schema constraints. Each scenario undergoes simulation rollouts, surrogate safety metric computation, and rule-based outcome labeling. To complement simulation-based evaluation, we introduce AgentDrive-MCQ, a 100,000-question multiple-choice benchmark spanning five reasoning dimensions: physics, policy, hybrid, scenario, and comparative reasoning. We conduct a large-scale evaluation of fifty leading LLMs on AgentDrive-MCQ. Results show that while proprietary frontier models perform best in contextual and policy reasoning, advanced open models are rapidly closing the gap in structured and physics-grounded reasoning. We release the AgentDrive dataset, AgentDrive-MCQ benchmark, evaluation code, and related materials at https://github.com/maferrag/AgentDrive

自主系统大模型评估自动驾驶场景生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。