让大模型通过显式状态追踪提升心智理论推理能力
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking

- 用PDDL语言显式表达环境状态与动作,解耦状态演化与信念推断
- 在多个心智理论基准上比现有最佳方法提升超5%准确率
- 适合研究大模型推理机制或需要可靠状态跟踪的场景
大语言模型在现有心智理论(ToM)基准上的表现远低于人类水平,即使使用思维链提示或概率信念更新也未能显著改善。我们认为,主要问题在于隐式状态跟踪不可靠,而非高层推理能力不足。为此提出PDDL-Mind,一种神经符号框架,将叙事描述转化为以规划领域定义语言(PDDL)表达的显式状态与动作,并通过预设领域验证动作引发的状态转移。该框架为大模型提供逻辑一致、显式的世界状态表示,用于心智理论任务。在MMToM-QA、MuMA和FanToM上的实验表明,PDDL-Mind在ToM问答任务上相较现有最优方法实现超过5%的绝对准确率提升。
原文摘要 · Abstract (English)
Large language models (LLMs) perform substantially below human level on existing theory-of-mind (ToM) benchmarks, even when augmented with chain-of-thought prompting or probabilistic belief updates. We argue that these failures primarily arise from unreliable implicit state tracking rather than limitations in high-level reasoning. We introduce PDDL-Mind, a neuro-symbolic framework that decouples environment state evolution from belief inference. By translating narrative descriptions into explicit states and actions expressed in Planning Domain Definition Language (PDDL), and by verifying action-induced state transitions against a predefined domain, PDDL-Mind provides LLMs with a logically consistent and explicit representation of world states for ToM tasks. Experiments on MMToM-QA, MuMA and FanToM show that PDDL-Mind achieves over 5% absolute accuracy gain over the best existing state-of-the-art method on ToM benchmark questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。