评测大模型对人类心理状态演变的追踪能力,发现其表现远逊于人类。
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
- 构建动态心理状态评估基准DynToM,涵盖1100个社交场景
- 十款主流大模型平均比人类低44.7%,状态演变推理能力弱
- 适合关注大模型社会认知能力的研究者与开发者
随着大型语言模型(LLMs)在人机交互中日益普及,评估其心智理论(ToM)能力——尤其是对动态心理状态的追踪能力——变得至关重要。现有评估基准多聚焦静态心理状态,忽略了真实社交互动中的时间演进特性。本文提出 extsc{DynToM},一个专为评估LLMs理解并跟踪跨情境心理状态演变能力而设计的新基准。通过系统化的四步框架,我们生成了1,100个社交情境,包含5,500个场景和78,100个问题,所有内容均经过真实性与质量验证。对十款先进大模型的全面评估显示,其平均表现比人类低44.7%,在追踪和推理心理状态转变时性能显著下降。这一差距凸显了当前大模型在建模人类心理动态性方面的根本局限。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present \textsc{DynToM}, a novel benchmark specifically designed to evaluate LLMs' ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7\%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs' ability to model the dynamic nature of human mental states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。