arXiv:2510.21329cs.CLcs.AI2025-10ACL被引 5

首个评估大模型应对旅行突发状况能力的基准测试

TripTide: A Benchmark for Adaptive Travel Planning under Disruptions

  • 构建包含多种扰动场景的旅行计划动态调整评测框架
  • 发现长行程计划地理一致性更好但应对突发情况能力下降
  • 适合研究智能旅行助手、鲁棒性评估的开发者和学者

近期工作如TripCraft和TravelPlanner推动了大语言模型(LLMs)在个性化、约束感知旅行行程生成中的应用。然而真实旅行常遭遇中断。为此,我们提出TripTide,首个评估LLMs在真实扰动下修订行程能力的基准。TripTide建模了扰动严重程度与旅客容忍度等关键维度,支持对航班取消、天气关闭或景点超售等事件的细致评估。我们开展三重验证:首先引入自动指标,包括意图保持度(修订计划的可行性和目标维持)、响应性(处理扰动的及时性与恰当性)以及适应性(原计划与修订计划在语义、空间和时序上的差异);其次采用大模型作为评判者进行自动质量评估;最后通过专家人工评估验证修订在语义、空间、时序和响应性方面的保留程度。实验表明,LLMs在时序一致性与语义稳定性上表现良好,短途行程的空间偏差较大但随行程延长而减小,说明长计划更易保持地理连贯性;然而随着计划长度增加,应对扰动的能力下降,揭示了大模型鲁棒性的局限。TripTide为评估基于大模型的旅行规划在现实不确定性下的自适应性、个性化与韧性建立了基准。

原文摘要 · Abstract (English)

Recent efforts like TripCraft and TravelPlanner have advanced the use of Large Language Models ( LLMs) for personalized, constraint aware travel itinerary generation. Yet, real travel often faces disruptions. To address this, we present TripTide, the first benchmark evaluating LLM's ability to revise itineraries under realistic disruptions. TripTide models key dimensions such as disruption severity and traveler tolerance, enabling nuanced assessment of LLM adaptability to events like flight cancellations, weather closures, or overbooked attractions. We conduct a threefold evaluation. First, we introduce automatic metrics including Preservation of Intent (how well the revised plan maintains feasibility and goals), Responsiveness (promptness and appropriateness of disruption handling), and Adaptability (semantic, spatial, and sequential divergence between original and revised plans). Second, we apply an LLM-as-a-judge approach to automatically assess revision quality. Third, we perform manual expert evaluation to verify whether revisions preserve semantic, spatial, sequential, and responsive aspects. Our experiments show that LLMs maintain strong sequential consistency and semantic stability, while spatial deviations are larger for shorter trips but decrease with longer ones, indicating that extended plans encourage better geographic coherence. However, disruption-handling ability declines as plan length increases, highlighting limits in LLM robustness. TripTide establishes a benchmark for evaluating adaptability, personalization, and resilience in LLM-based travel planning under real-world uncertainty.

旅行规划大模型评估鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。