构建真实旅行规划评测基准,测试大模型多轮交互与工具使用能力
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
- 基于真实场景构建多轮旅行规划任务,涵盖需求挖掘与边界识别
- 引入10个真实旅行工具缓存结果,支持稳定可复现的评估
- 发现顶尖模型在不同能力上表现不均,揭示其真实应用短板
旅行规划是检验大语言模型(LLMs)规划与工具使用能力的自然现实任务。现有研究仍与真实需求存在差距,主要源于领域覆盖有限、多轮对话中用户隐式偏好的建模不足,以及对智能体能力边界的评估缺失。为此,我们提出《TravelBench》,一个面向真正真实世界旅行规划的评测基准。通过收集真实场景中的用户查询、偏好和工具,构建三个子任务:单轮、多轮和不可解任务,以评估智能体在三种核心能力上的表现:(1)独立解决问题,(2)与用户交互以挖掘隐式偏好,(3)识别自身能力边界。为实现稳定的工具调用和可复现的评估,我们缓存真实工具调用结果,并搭建集成十种旅行相关工具的沙盒环境,使智能体能组合这些工具解决大多数实际旅行规划问题。我们在TravelBench上评估多个LLM,发现即使先进模型在不同能力间也表现出明显不平衡。进一步系统验证表明该基准具有稳定性。TravelBench为推进真实世界旅行规划的LLM智能体研究提供了实用且可复现的评测基础。
原文摘要 · Abstract (English)
Travel planning is a natural real-world task to test large language models' (LLMs) planning and tool-use abilities. Although prior work has studied LLM performance on travel planning, existing settings still differ from real-world needs, mainly due to limited domain coverage, insufficient modeling of users' implicit preferences in multi-turn conversations, and a lack of evaluation of agents' capability boundaries. To mitigate these gaps, we propose $\textbf{TravelBench}$, a benchmark for $\textit{truly real-world}$ travel planning. We collect user queries, user preferences, and tools from real scenarios, and construct three subtasks -- $\textit{Single-Turn}$, $\textit{Multi-Turn}$, and $\textit{Unsolvable}$ -- to evaluate agents' three core capabilities in real settings: (1) solving problems independently, (2) interacting with users to elicit implicit preferences, and (3) recognizing the capability boundaries. To enable stable tool invocation and reproducible evaluation, we cache real tool-call results and build a sandbox environment which integrates ten travel-related tools, enabling agents to combine these tools to solve most practical travel planning problems. We evaluate multiple LLMs on TravelBench and find that even advanced models exhibit imbalanced performance across different capabilities. Our further systematic verification demonstrates the stability of the proposed benchmark. TravelBench provides a practical and reproducible benchmark to advance research on LLM agents for real-world travel planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。