构建细粒度旅行规划评估体系,用统一奖励提升真实场景规划质量
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
- 设计统一奖励机制,融合可行性、可靠性与吸引力等细粒度指标
- 在4870条真实查询上测试,强化学习显著提升行程可行性与综合评分
- 适合研究旅行规划、RL应用或人机交互的开发者与研究人员
旅行规划是一项重要但复杂的任务,即使对先进的大语言模型(LLM)也构成挑战。现有基准在评估行程可行性、可靠性和用户参与度方面存在不足。本文提出一个综合性旅行规划评估基准,将细粒度评价标准整合为单一奖励信号,支持直接比较方案质量并无缝接入强化学习(RL)。评估器与旅行专家标注达到60.75%的中等一致性,优于多个基于LLM的评判基线。我们还发布了包含4870个查询的大规模数据集,其中219个为真实世界的自由格式请求,以促进对用户真实意图的泛化。通过该基准,我们在多种方法和大模型上进行了广泛实验,包括测试时计算、神经符号方法、监督微调及基于GRPO的强化学习。结果表明,在基础模型上,强化学习普遍优于仅用提示和监督微调的方法,显著提升行程可行性和统一奖励得分。
原文摘要 · Abstract (English)
Travel planning is a valuable yet complex task that poses significant challenges even for advanced large language models (LLMs). While recent benchmarks have advanced in evaluating LLMs' planning capabilities, they often fall short in evaluating feasibility, reliability, and engagement of travel plans. We introduce a comprehensive benchmark for travel planning that unifies fine-grained criteria into a single reward, enabling direct comparison of plan quality and seamless integration with reinforcement learning (RL). Our evaluator achieves moderate agreement with travel-expert annotations (60.75%) and outperforms multiple LLM-as-judge baselines. We further release a large-scale dataset of 4,870 queries including 219 real-world, free-form requests for generalization to authentic user intent. Using this benchmark, we conduct extensive experiments across diverse methods and LLMs, including test-time computation, neuro-symbolic approaches, supervised fine-tuning, and RL via GRPO. Across base models, RL generally improves itinerary feasibility over prompt-only and supervised baselines, yielding higher unified reward scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。