arXiv:2410.13648cs.CLcs.AI2024-10被引 51

测试大模型在真实场景中能否把心理理解转化为行为预测。

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

  • 设计新基准,分层测试心理状态推理与实际行为应用能力
  • 模型能准确判断他人是否知情,但难以预测其后续行为
  • 揭示大模型社会推理的隐性短板,适合研究人机交互与认知评估

大型语言模型(LLMs)正被越来越多地测试其‘心智理论’(ToM)能力——即理解自己与他人心理状态的能力。然而,现有评估多停留在经典玩具故事中的显式信念推断,未能检验模型是否能在多样日常场景中隐式运用此类知识来预测或评判人类行为。本文提出SimpleToM,一个面向多重层次ToM推理的新基准:从心理状态推断(显式ToM)到行为预测与判断(应用ToM)。任务设定于超市、医院、学校、办公室等真实场景,信息不对称自然存在(如商品霉变、医患信息不全、设备锁定等)。每个短故事配三个问题,分别考察:(a) 心理状态(玛丽是否知道发霉?)、(b) 行为预测(她会付款还是报告?)、(c) 判断(她付款合理吗?)。实验显示,顶尖模型虽能可靠推断心理状态(a),但在行为预测(b)和判断(c)上表现急剧下降,暴露出其在知识应用层面的显著脆弱性,凸显了‘已知’与‘可用’之间的巨大鸿沟。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly tested for a "Theory of Mind" (ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at explicit belief attribution in classical toy stories or stylized tasks, leaving open the questions of whether LLMs can implicitly apply such knowledge to predict human behavior, or to judge an observed behavior, in diverse scenarios. We introduce SimpleToM, a benchmark that advances ToM evaluation along two novel axes. First, it probes multiple levels of ToM reasoning, from mental state inference (explicit ToM) to behavior prediction and judgment (applied ToM). Second, it situates these tasks in diverse, everyday scenarios - such as supermarkets, hospitals, schools, and offices - where information asymmetries naturally arise (e.g., hidden defects in grocery store items, incomplete information in provider-patient interactions, or restricted access to locked devices). SimpleToM contains concise stories (e.g., "The can of Pringles has moldy chips in it. Mary picks up the can in the supermarket and walks to the cashier."), each with three questions that test different degrees of ToM reasoning, asking models to predict: (a) mental states ("Is Mary aware of the mold?"), (b) behaviors ("Will Mary pay for the chips or report the mold?"), and (c) judgments ("Mary paid for the chips. Was that reasonable?"). Experiments reveal a striking gap: state-of-the-art models often reliably infer mental state (a), but fail at applying knowledge about the mental state for secondary predictions, with performance dropping sharply for behavior prediction (b) and further for behavior judgment (c). This exposes a critical fragility in LLMs' social reasoning in terms of what they know (explicit ToM) versus how well they can implicitly apply that knowledge for predictions (applied ToM).

心智理论大模型评测社会推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。