arXiv:2605.15333cs.AI2026-05

用大模型零样本识别目标,检验其真实推理能力

Zero-Shot Goal Recognition with Large Language Models

  • 让大模型在无训练情况下判断目标状态,基于世界知识推理
  • 部分模型随证据增加准确率提升,接近基于地标的方法
  • 适合评估大模型的规划认知基础,揭示推理机制差异

大语言模型在经典规划领域已接近传统规划器表现,但依赖世界知识而非真正的符号推理。目标识别是一种互补的归纳任务,更契合大模型优势:只需评估与世界知识的一致性,无需生成新动作序列。本文首次系统地在关键PDDL基准上对前沿大模型进行零样本目标识别评估。结果表明,大模型在目标识别上的表现参差不齐:部分模型随证据积累准确率提升,全观测下接近基于地标的方法;而另一些模型始终受世界知识先验束缚,无法吸收更多证据。定性分析显示,这种差异源于证据整合机制的根本不同,而非领域熟悉度。该研究将目标识别定位为评估大模型基础规划知识的可靠基准。

原文摘要 · Abstract (English)

Large language models have recently reached near-parity with classical planners on well-known planning domains, yet this competence relies on world-knowledge exploitation rather than genuine symbolic reasoning. Goal recognition is a complementary abductive task structurally better suited to LLM strengths: it consists of evaluating consistency with world knowledge rather than generating novel action sequences. This paper provides the first systematic zero-shot evaluation of frontier LLMs as goal recognisers on key classical PDDL benchmarks. Our results show that LLM competence on goal recognition is uneven: some models scale with evidence and approach landmark-based accuracy at full observations, while others remain anchored to world-knowledge priors regardless of how much evidence accumulates. Qualitative analysis of model reasoning traces reveals that this divergence reflects a fundamental difference in evidence integration rather than domain familiarity. These findings position goal recognition as a principled benchmark for the foundational planning knowledge of LLMs.

大模型推理目标识别规划能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。