arXiv:2605.03308cs.AI2026-05

拆解旅行规划能力,发现大模型在隐含需求理解上严重不足

Revisiting the Travel Planning Capabilities of Large Language Models

论文配图:Revisiting the Travel Planning Capabilities of Large Language Models
图 1 · 摘自论文原文
  • 将旅行规划分解为五项基础能力,逐项评估性能边界
  • 模型能提取明确约束,却难以识别开放世界的隐含需求
  • 适合研究大模型长程推理缺陷与改进方向的学者参考

旅行规划是长时序推理的关键任务,暴露出大语言模型(LLMs)的重大短板。然而,现有基准多以端到端方式评估最终计划,缺乏可解释性,难以分析失败根源。为此,我们将旅行规划拆解为五个原子子能力:约束提取(Constraint Extraction)、工具使用(Tool Use)、计划生成(Plan Generation)、错误识别(Error Identification)和错误纠正(Error Correction)。通过引入理想中间上下文的解耦评估协议,严格隔离各组件,测量其原子性能上限,避免级联错误干扰。结果揭示显著差异:尽管模型在提取显式约束方面表现良好,但在推断隐含的开放世界需求上存在明显不足;此外,计划生成中存在结构性偏差,自我纠错能力差,表现为过度敏感且错误坚持。这些发现为提升大模型推理与规划能力提供了精确方向。

原文摘要 · Abstract (English)

Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluations primarily assess final plans in an end-to-end manner, which lacks interpretability and makes it difficult to analyze the root causes of failures. To bridge this gap, we decompose travel planning into five constituent atomic sub-capabilities, including \emph{Constraint Extraction}, \emph{Tool Use}, \emph{Plan Generation}, \emph{Error Identification}, and \emph{Error Correction}. We implement a decoupled evaluation protocol leveraging oracle intermediate contexts to rigorously isolate these components, thereby measuring the atomic performance boundary without the noise of cascading errors. Our results highlight a clear contrast in performance: while LLMs are proficient in extracting explicit constraints, they struggle to infer implicit, open-world requirements. Furthermore, they exhibit structural biases in plan generation and suffer from ineffective self-correction, characterized by excessive sensitivity and erroneous persistence. These findings offer precise directions for improving LLM reasoning and planning abilities.

大模型推理旅行规划能力评估长程思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。