现有大模型心理理论测试失效,因无法评估其对新伙伴的适应能力。
Position: Theory of Mind Benchmarks are Broken for Large Language Models
- 提出功能心理理论:衡量模型根据伙伴行为理性调整自身策略的能力。
- 发现开源大模型虽有强预测力,却难适应简单伙伴策略,暴露适应性短板。
- 强调长期互动中的动态适应是真实心理理论能力的核心,适合评估大模型认知韧性。
本文指出,当前多数大语言模型(LLM)的心理理论评测体系存在根本缺陷,因其无法直接检验模型在面对新伙伴时的适应能力。这一问题源于人类心理理论测试方法的过度模仿,错误地将人类一致性推理预期投射到AI上。事实上,现有评测仅测量‘字面心理理论’——即预测他人行为的能力,而此指标仅在代理具备自我一致性推理时才有效。为此,我们提出‘功能心理理论’:模型在上下文内根据伙伴行为做出合理响应的能力。实验表明,许多开源大模型虽展现出强大的字面心理理论能力,但在应对简单伙伴策略时,尤其在长交互序列中,表现明显不足。这说明字面能力与功能适应能力不可互换。实现真正的功能心理理论,尤其是在长期互动中,是未来评估大模型认知成熟度的关键挑战。
原文摘要 · Abstract (English)
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks for LLMs are overwhelmingly inspired by the methods used to test theory of mind in humans and fall victim to a fallacy of attributing human-like qualities to AI agents. We expect that humans will engage in a consistent reasoning process across various questions about a situation, but this is known to not be the case for current LLMs. Most theory of mind benchmarks only measure what we call literal theory of mind: the ability to predict the behavior of others. However, this type of metric is only informative when agents exhibit self-consistent reasoning. Thus, we introduce the concept of functional theory of mind: the ability to adapt to agents in-context following a rational response to their behavior. We find that many open source LLMs are capable of displaying strong literal theory of mind capabilities, but seem to struggle with functional theory of mind -- even with exceedingly simple partner policies. Simply put, strong literal theory of mind performance does not necessarily imply strong functional theory of mind performance or vice versa. Achieving functional theory of mind, particularly over long interaction horizons with a partner, is a significant challenge deserving a prominent role in any meaningful LLM theory of mind evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。