用神经知识库提升大模型的心智推理能力,尤其擅长高阶推理。
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
- 引入实体状态神经知识库,通过迭代掩码实现精准视角转换。
- 在多个基准上显著提升各尺寸模型的高阶心智推理表现。
- 适合研究大模型社会认知与复杂推理的学者使用。
心智理论(ToM)是理解他人感知与心理状态的能力,对人类交互至关重要,但对大语言模型仍具挑战。现有方法依赖现成大模型进行视角推断,效率低且难以支持高阶推理。为此,我们提出EnigmaToM,一种新型神经符号框架,通过集成实体状态神经知识库(Enigma),实现(1)心理学启发的迭代掩码机制以促进准确视角转换,(2)关键实体信息注入。Enigma生成结构化实体状态知识,构建空间场景图以追踪信念变化,并丰富事件中的细粒度状态细节。在ToMi、HiToM和FANToM基准上的实验表明,EnigmaToM显著提升不同规模大模型的ToM推理能力,尤其在高阶推理场景中表现优异。
原文摘要 · Abstract (English)
Theory-of-Mind (ToM), the ability to infer others' perceptions and mental states, is fundamental to human interaction but remains challenging for Large Language Models (LLMs). While existing ToM reasoning methods show promise with reasoning via perceptual perspective-taking, they often rely excessively on off-the-shelf LLMs, reducing their efficiency and limiting their applicability to high-order ToM reasoning. To address these issues, we present EnigmaToM, a novel neuro-symbolic framework that enhances ToM reasoning by integrating a Neural Knowledge Base of entity states (Enigma) for (1) a psychology-inspired iterative masking mechanism that facilitates accurate perspective-taking and (2) knowledge injection that elicits key entity information. Enigma generates structured knowledge of entity states to build spatial scene graphs for belief tracking across various ToM orders and enrich events with fine-grained entity state details. Experimental results on ToMi, HiToM, and FANToM benchmarks show that EnigmaToM significantly improves ToM reasoning across LLMs of varying sizes, particularly excelling in high-order reasoning scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。