构建真实旅行规划评估框架,全面检验大模型的多维决策能力。
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

- 六维评估体系覆盖时空成本、预算、合规等维度
- 基于真实住宿价与交通数据模拟完整行程
- 适合研究智能旅行代理与多目标规划的学者
大语言模型(LLMs)推动了旅行规划应用的发展,但现有评估基准存在三大局限:过度强调约束符合性,忽视时空成本等多维质量;数据集缺乏真实性和关键领域(如住宿、交通)覆盖;孤立评估每日计划,忽略住宿安排与行程节奏对整体计划的影响。为此,我们提出TravelEval,一个真实且全面的评估框架。TravelEval包含:1)新颖的六维评估体系,涵盖准确性、合规性、时间性、空间性、经济性与实用性;2)高保真的数据沙盒,包含精确的住宿价格与真实的城际交通数据;3)基于模拟的全局评估方法,集成地理信息与细粒度排队时间,模拟完整旅行计划。使用TravelEval评估12种主流方法发现:LLMs在全局优化的多维规划中表现不佳(尤其在时空推理与预算合规方面),且代理式推理策略未带来稳定提升。综上,TravelEval通过基于时空真实模拟与综合指标,为大模型驱动的旅行规划研究与应用提供坚实评估基础。
原文摘要 · Abstract (English)
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.g., lodging, transport); and 3) isolated daily plan assessments that miss critical details (e.g., the impact of daily accommodation and visit pacing) needed for entire plan's evaluation. To address this gap, we introduce TravelEval, a realistic and comprehensive benchmark. TravelEval features 1) a novel six-dimensional evaluation framework to holistically assess plans across accuracy, compliance, temporality, spatiality, economy, and utility dimensions; 2) a highly realistic data sandbox with precise accommodation pricing and authentic intercity transportation data; and 3) a simulation-based global evaluation method that emulates complete travel plans with API-integrated geographic information and fine-grained queuing time. Evaluating 12 mainstream approaches with TravelEval reveals several valuable insights, such that LLMs struggle with globally-optimized multi-dimensional planning (especially in spatio-temporal reasoning and budget compliance), and agentic reasoning strategies offer no consistent improvement. Concisely, TravelEval facilitates travel plan evaluation via grounded spatio-temporal emulation and comprehensive metrics, providing a robust foundation for advancing LLM-powered travel planning research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。