arXiv:2601.04698cs.AIcs.CL2026-01被引 4

多路径推理+约束门控强化学习,提升旅行规划的可行性和用户偏好匹配度。

TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning

  • 采用多路径思维链探索更广解空间,避免单一路径局限。
  • 在基准测试中显著超越现有方法,提升可行性和用户偏好对齐。
  • 动态优先满足硬约束后才优化软约束,适合复杂旅行规划场景。

旅行规划是需综合多维度信息的复杂决策过程。现有方法面临三大挑战:(1)候选兴趣点(POI)筛选时难以兼顾高召回率;(2)单一推理路径限制了可行解空间的探索能力;(3)同时优化硬约束与软约束仍具难度。为此,我们提出TourPlanner框架,融合多路径推理与约束门控强化学习。首先,引入个性化召回与空间优化(PReSO)流程构建空间感知的候选POI集合;随后,提出竞争性共识思维链(CCoT),通过多路径推理增强解空间探索能力;最后,在强化学习阶段集成基于Sigmoid的门控机制,仅在硬约束满足后动态优先优化软约束。在旅行规划基准上的实验表明,TourPlanner达到当前最优性能,显著优于现有方法,在可行性与用户偏好对齐方面均有提升。

原文摘要 · Abstract (English)

Travel planning is a sophisticated decision-making process that requires synthesizing multifaceted information to construct itineraries. However, existing travel planning approaches face several challenges: (1) Pruning candidate points of interest (POIs) while maintaining a high recall rate; (2) A single reasoning path restricts the exploration capability within the feasible solution space for travel planning; (3) Simultaneously optimizing hard constraints and soft constraints remains a significant difficulty. To address these challenges, we propose TourPlanner, a comprehensive framework featuring multi-path reasoning and constraint-gated reinforcement learning. Specifically, we first introduce a Personalized Recall and Spatial Optimization (PReSO) workflow to construct spatially-aware candidate POIs' set. Subsequently, we propose Competitive consensus Chain-of-Thought (CCoT), a multi-path reasoning paradigm that improves the ability of exploring the feasible solution space. To further refine the plan, we integrate a sigmoid-based gating mechanism into the reinforcement learning stage, which dynamically prioritizes soft-constraint satisfaction only after hard constraints are met. Experimental results on travel planning benchmarks demonstrate that TourPlanner achieves state-of-the-art performance, significantly surpassing existing methods in both feasibility and user-preference alignment.

旅行规划多路径推理强化学习约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。