arXiv:2606.21169cs.AI2026-06被引 1

评测智能体在个性化旅行规划中的综合能力,发现现有模型常忽视用户疲劳。

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

论文配图:Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
图 1 · 摘自论文原文
  • 构建动态交互的分钟级行程规划任务,结合用户画像与突发状况。
  • 18个语言模型表现差距明显,多数生成行程虽可行却极度耗时耗力。
  • 首次用大模型模拟器评估疲劳等主观体验,适合旅行规划研究者参考。

交互式旅行规划已成为语言模型的重要应用。智能体需在多轮对话中应对偏好变化和意外干扰,做出复杂的、基于个人画像的决策。然而,现有基准往往孤立评估可行性、个性化或交互能力。为此,我们提出Trip+,用于衡量智能体进行整体旅行规划的能力。在Trip+中,给定旅行者画像和动态交互,智能体需生成并修订分钟级行程。通过基于大模型的模拟器评估端到端旅行体验,可量化疲劳等主观指标。场景涵盖从简单请求解决到环境驱动的复杂重规划。我们评估了18个语言模型,发现体验质量存在持续差距:模型倾向于生成技术上可行但极其耗能的行程,严重偏离用户画像偏好。

原文摘要 · Abstract (English)

Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make complex, profile-conditioned planning decisions. However, existing benchmarks often evaluate feasibility, personalization, or interaction in relatively isolated settings. We therefore introduce Trip+ to measure the ability of agents to plan travel holistically. In Trip+, given traveler profiles and dynamic interactions, agents must generate and revise minute-level itineraries. End-to-end traveler experiences are evaluated via an LLM-based simulator, enabling the assessment of subjective metrics like fatigue. Our scenarios range from simple request resolutions to complex environment-driven replanning. We evaluate 18 LMs and find a consistent gap in experiential quality. Models favor technically feasible but exhausting itineraries that diverge sharply from profiled traveler preferences.

旅行规划智能体用户体验评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。