arXiv:2602.22638cs.AI2026-02KDD被引 8

构建真实出行场景评估基准,测试大模型路线规划能力

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

  • 基于高德真实用户查询构建多城市路线规划数据集
  • 设计确定性沙箱环境,确保评估结果可复现
  • 发现当前模型在个性化路线规划上表现不足

基于大语言模型的路线规划代理通过自然语言交互和工具协同决策,为日常出行提供新范式。然而,真实出行场景下的系统化评估受限于多样化的导航需求、非确定性地图服务及可复现性差等问题。本文提出MobilityBench,一个可扩展的真实出行场景评估基准。该基准源自高德平台大规模匿名用户查询,覆盖全球多个城市的多样化路线规划意图。为实现可复现的端到端评估,我们设计了确定性API回放沙箱,消除实时服务带来的环境差异。进一步提出以结果有效性为核心的多维度评估协议,涵盖指令理解、规划能力、工具使用与效率。利用MobilityBench,我们对多种基于LLM的路线规划代理进行跨场景评估,并深入分析其行为表现。结果显示,现有模型在基础信息检索与路线规划任务中表现良好,但在偏好约束型路线规划任务中明显乏力,表明个性化出行应用仍有显著提升空间。相关数据、评估工具包及文档已开源至https://github.com/AMAP-ML/MobilityBench。

原文摘要 · Abstract (English)

Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through natural language interaction and tool-mediated decision making. However, systematic evaluation in real-world mobility settings is hindered by diverse routing demands, non-deterministic mapping services, and limited reproducibility. In this study, we introduce MobilityBench, a scalable benchmark for evaluating LLM-based route-planning agents in real-world mobility scenarios. MobilityBench is constructed from large-scale, anonymized real user queries collected from Amap and covers a broad spectrum of route-planning intents across multiple cities worldwide. To enable reproducible, end-to-end evaluation, we design a deterministic API-replay sandbox that eliminates environmental variance from live services. We further propose a multi-dimensional evaluation protocol centered on outcome validity, complemented by assessments of instruction understanding, planning, tool use, and efficiency. Using MobilityBench, we evaluate multiple LLM-based route-planning agents across diverse real-world mobility scenarios and provide an in-depth analysis of their behaviors and performance. Our findings reveal that current models perform competently on Basic information retrieval and Route Planning tasks, yet struggle considerably with Preference-Constrained Route Planning, underscoring significant room for improvement in personalized mobility applications. We publicly release the benchmark data, evaluation toolkit, and documentation at https://github.com/AMAP-ML/MobilityBench.

路线规划大模型评估基准出行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。