构建真实旅行场景长对话基准,测试大模型多工具协作与长期适应能力。
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
- 基于真实旅行数据设计多轮交互任务,支持150+工具调用与200k+上下文长度。
- 顶尖模型在困难子集上成功率不足10%,凸显长时交互挑战。
- 适合研究长期对话、多工具协同与在线强化学习的学者与工程师。
随着基于大语言模型的智能体被部署于日益复杂的现实场景中,现有基准未能充分反映全局约束执行、多工具推理协调以及长期多轮交互中用户行为演变等关键挑战。为此,我们提出 extbf{TRIP-Bench},一个基于真实旅行规划场景的长时交互基准。该基准利用真实世界数据,提供18个精选工具和40多个旅行需求,支持自动化评估,并包含不同难度划分。其中,困难子集强调长时间、模糊性交互、风格变化、可行性变动及迭代版本修改。对话可长达15轮用户输入,涉及超过150次工具调用,上下文长度超过20万词符。实验表明,即使先进模型在简单子集上成功率也仅达50%,在困难子集上更低至10%以下。我们进一步提出 extbf{GTPO},一种具备特殊奖励归一化与差分奖励的在线多轮强化学习方法。将其应用于 Qwen2.5-32B-Instruct 后,显著提升约束满足率与交互鲁棒性,在我们的评估中优于 Gemini-3-Pro。我们期望 TRIP-Bench 能推动实用长时交互智能体的发展,而 GTPO 可为鲁棒长时训练提供有效在线强化学习方案。
原文摘要 · Abstract (English)
As LLM-based agents are deployed in increasingly complex real-world settings, existing benchmarks underrepresent key challenges such as enforcing global constraints, coordinating multi-tool reasoning, and adapting to evolving user behavior over long, multi-turn interactions. To bridge this gap, we introduce \textbf{TRIP-Bench}, a long-horizon benchmark grounded in realistic travel-planning scenarios. TRIP-Bench leverages real-world data, offers 18 curated tools and 40+ travel requirements, and supports automated evaluation. It includes splits of varying difficulty; the hard split emphasizes long and ambiguous interactions, style shifts, feasibility changes, and iterative version revision. Dialogues span up to 15 user turns, can involve 150+ tool calls, and may exceed 200k tokens of context. Experiments show that even advanced models achieve at most 50\% success on the easy split, with performance dropping below 10\% on hard subsets. We further propose \textbf{GTPO}, an online multi-turn reinforcement learning method with specialized reward normalization and reward differencing. Applied to Qwen2.5-32B-Instruct, GTPO improves constraint satisfaction and interaction robustness, outperforming Gemini-3-Pro in our evaluation. We expect TRIP-Bench to advance practical long-horizon interactive agents, and GTPO to provide an effective online RL recipe for robust long-horizon training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。