首个面向多人旅行规划的基准,测试大模型在多轮对话中协调偏好与冲突的能力。
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

- 构建真实用户数据驱动的多人多轮旅行规划任务
- 现有最强模型计划有效率不足12%,公平性与协调能力严重不足
- 聚焦偏好挖掘、冲突调和与公平性平衡,填补领域空白
现实中的旅行规划本质上是群体活动,但现有大模型旅行规划基准大多简化为单人任务,导致研究趋于饱和。该单人假设忽略了群体规划的核心难点:跨用户的私有偏好发现、冲突识别以及效用与公平的权衡。为此,我们提出首个面向多用户、多轮对话的旅行规划基准——GroupTravelBench,基于真实用户画像、兴趣点数据及票务价格构建,包含650个任务,分为三个难度等级。每个任务在同步群聊沙盒中运行,并配备缓存工具数据以支持可复现的离线评估。除常规多步推理与工具使用外,该基准还重点考察三项群体专属能力:(i) 通过多轮对话挖掘私有偏好;(ii) 通过妥协或分组解决用户间冲突;(iii) 平衡群体整体效用与公平性。我们配套设计了结合规则化结果指标与大模型判断过程指标的评估框架。在多个前沿模型上测试发现,即使最强模型在四项规则化结果指标上均未达标,计划有效性低于12%,表明群体层面的规划质量仍是大模型旅行代理的关键挑战。
原文摘要 · Abstract (English)
Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approaching saturation. This single-user assumption sidesteps what makes group planning hard for an agent: discovering private preferences across multiple users, surfacing conflicts, and balancing utility against fairness. To bring the task back to its multi-user reality, we introduce \textbf{\textit{GroupTravelBench}}, the first benchmark for \textbf{multi-user, multi-turn} travel planning. Built from real user profiles, POI data, and ticket prices, it comprises 650 tasks across three difficulty levels, each running in a synchronous group-chat sandbox with cached tool data for reproducible offline evaluation. Beyond the multi-step reasoning and tool use that single-user benchmarks already test, GroupTravelBench probes three group-specific capabilities: \textit{(i) elicitation} of private preferences through multi-turn dialogue; \textit{(ii) coordination} of inter-user conflicts via compromise or subgrouping; and \textit{(iii) planning} that balances group utility against fairness. We pair this with a complementary evaluation framework combining rule-based outcome metrics and LLM-judge process metrics. Across a wide range of frontier models, even the strongest agents fall short on all four rule-based outcome metrics, with plan validity below 12\%, suggesting that group-level outcome quality is a key open challenge for LLM travel-planning agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。