为旅行规划引入不确定性评估,测试模型在延误和人流变化下的应对能力。
UTP-Bench: Uncertainty-aware Travel Planning Benchmark

- 基于印度504个城市的真实数据,构建含交通延迟与人流波动的动态旅行数据集。
- 提出三项新指标,量化行程对突发延误和人流变化的抗干扰能力。
- 发现当前大模型生成的行程在时间缓冲和人群敏感调度上远不如人工设计。
大型语言模型(LLMs)在自动化旅行路线生成方面展现出强大能力。然而,现实旅行规划本质上具有不确定性:交通延误、人流波动及意外随机延迟常导致原本可行的行程失效。现有基准如TravelPlanner和TripCraft假设确定性环境,仅评估静态约束满足性,忽略生成计划在不确定性下的鲁棒性。为此,我们提出UTP-Bench,一个大规模的不确定性感知旅行规划基准。该数据集整合了覆盖印度504个城市的实际旅行数据,包括景点、餐厅、住宿及多模式交通网络。为模拟真实中断,UTP-Bench融合了来自主要城市收集的实证延迟分布和人流密度模式,支持在随机条件下评估旅行计划。我们进一步提出三项评估指标:缓冲充足度评分(BAS)、人群感知时间评分(CATS)和交通延迟吸收评分(TDAS),用于量化生成行程对交通延误和人流变化的鲁棒性。对GPT-5、Qwen3、Mistral和Phi-4等先进大模型的实验表明,模型生成计划与人工设计计划之间存在显著差距,尤其体现在时间缓冲、延迟感知交通调度和人群敏感规划方面。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction and ignoring whether generated plans remain robust when such uncertainties arise. To address this limitation, we introduce UTP-Bench1 , a large-scale benchmark for uncertainty-aware travel planning. The dataset integrates real-world travel data spanning 504 cities of India, including attractions, restau- rants, accommodations, and multi-modal trans- portation networks. To model realistic disrup- tions, UTP-Bench incorporates empirical delay distributions and crowd-density patterns col- lected from major cities, enabling evaluation of travel plans under stochastic conditions. We further propose three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which quan- tify the ability of generated itineraries to main- tain robustness against transit delays and crowd variability. Experiments with state-of-the-art LLMs like GPT-5, Qwen3, Mistral and Phi-4 re- veal substantial gaps between model-generated and human-authored plans, particularly in tem- poral buffering, delay-aware transportation scheduling, and crowd-sensitive planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。