arXiv:2505.00026cs.CLcs.AI2025-05ACL被引 36

评估并提升大模型理解人类心理状态的能力

Theory of Mind in Large Language Models: Assessment and Enhancement

  • 基于故事类评测基准分析大模型的心理推理能力
  • 提出增强大模型心智理论能力的最新方法
  • 适合研究智能对话与社会认知的学者参考

心智理论(ToM)——即对自身及他人心理状态进行推理的能力——是人类社交智能的核心。随着大语言模型(LLMs)日益融入日常生活,理解其解读和回应人类心理状态的能力,对于实现有效交互至关重要。本文通过分析近期提出的广泛使用的故事类评测基准,系统回顾了大模型的ToM能力。同时,深入剖析了旨在提升大模型ToM能力的最新方法。此外,本文还指出了未来研究的有前景方向,以进一步推动该能力发展,并使大模型更好地适应更真实、多样的应用场景。本综述为关注评估与提升大模型心智理论能力的研究者提供了重要资源。

原文摘要 · Abstract (English)

Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become increasingly integrated into daily life, understanding their ability to interpret and respond to human mental states is crucial for enabling effective interactions. In this paper, we review LLMs' ToM capabilities by analyzing both evaluation benchmarks and enhancement strategies. For evaluation, we focus on recently proposed and widely used story-based benchmarks. For enhancement, we provide an in-depth analysis of recent methods aimed at improving LLMs' ToM abilities. Furthermore, we outline promising directions for future research to further advance these capabilities and better adapt LLMs to more realistic and diverse scenarios. Our survey serves as a valuable resource for researchers interested in evaluating and advancing LLMs' ToM capabilities.

心智理论大模型社交智能评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。