arXiv:2502.08796cs.CLcs.CY2025-02综述被引 14

系统梳理大模型在心理理论任务中的评估方法与局限。

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

  • 按认知科学分类框架整合现有评测基准与任务
  • 发现大模型虽有初步心理理论能力但远未达人类水平
  • 适合关注AI认知能力评估的研究者与伦理设计者

近年来,评估大型语言模型(LLMs)在心理理论(ToM)任务中的能力受到研究界广泛关注。随着领域快速发展,多样化的评估方法和路径日益复杂。本文通过基于认知科学的分类体系,系统综述了当前评估LLMs执行ToM任务的努力。尽管取得显著进展,但其在心理理论方面的表现仍存争议。该综述批判性分析了评估技术、提问策略及模型内在局限,揭示出虽然大模型在部分任务中展现出初步的心理状态推理能力,但在模仿人类认知方面仍存在显著差距。

原文摘要 · Abstract (English)

In recent years, evaluating the Theory of Mind (ToM) capabilities of large language models (LLMs) has received significant attention within the research community. As the field rapidly evolves, navigating the diverse approaches and methodologies has become increasingly complex. This systematic review synthesizes current efforts to assess LLMs' ability to perform ToM tasks, an essential aspect of human cognition involving the attribution of mental states to oneself and others. Despite notable advancements, the proficiency of LLMs in ToM remains a contentious issue. By categorizing benchmarks and tasks through a taxonomy rooted in cognitive science, this review critically examines evaluation techniques, prompting strategies, and the inherent limitations of LLMs in replicating human-like mental state reasoning. A recurring theme in the literature reveals that while LLMs demonstrate emerging competence in ToM tasks, significant gaps persist in their emulation of human cognitive abilities.

心理理论大模型评估认知能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。