用模拟和分解提升大模型的共情推理能力
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
- 通过模拟用户视角递归分解共情任务
- 在复杂共情任务上显著超越基线方法
- 无需训练,少量提示即可适配多场景
心智理论(ToM)是理解他人心理状态的能力。尽管这对人类互动至关重要,但测试显示大语言模型对此仅有初步理解。虽然最先进的闭源模型在某些任务上接近人类表现,但在涉及更复杂结构化推理的任务中仍表现不佳。本文受认知心理学中‘假装游戏’或‘模拟理论’启发,提出基于大模型的推理算法Decompose-ToM,通过递归模拟用户视角,将复杂共情任务分解为:主体识别、问题重构、世界模型更新和知识可用性判断等简化函数。我们在高阶共情任务及对话情境下的共情能力测试中验证该方法,结果表明其在不同模型上均显著优于基线方法,且仅需极少提示调优,无需额外模型训练。
原文摘要 · Abstract (English)
Theory of Mind (ToM) is the ability to understand and reflect on the mental states of others. Although this capability is crucial for human interaction, testing on Large Language Models (LLMs) reveals that they possess only a rudimentary understanding of it. Although the most capable closed-source LLMs have come close to human performance on some ToM tasks, they still perform poorly on complex variations of the task that involve more structured reasoning. In this work, we utilize the concept of "pretend-play", or ``Simulation Theory'' from cognitive psychology to propose ``Decompose-ToM'': an LLM-based inference algorithm that improves model performance on complex ToM tasks. We recursively simulate user perspectives and decompose the ToM task into a simpler set of functions: subject identification, question-reframing, world model updation, and knowledge availability. We test the algorithm on higher-order ToM tasks and a task testing for ToM capabilities in a conversational setting, demonstrating that our approach shows significant improvement across models compared to baseline methods while requiring minimal prompt tuning across tasks and no additional model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。