arXiv:2412.13631cs.AIcs.CL2024-12ACL被引 13

提出评估大模型心智理论应关注推理深度而非仅逻辑正确性

Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning

  • 区分心智理论的两步:判断是否需要深层推理,再执行对应推断
  • 指出当前研究多聚焦静态逻辑问题,忽略动态情境下的推理深度
  • 建议借鉴认知科学实验设计,提升对心智理论能力的评估效度

大语言模型的心智理论(ToM)能力成为近期研究热点。认知科学将ToM任务分为两步:1)判断是否启用心智理论及所需推理深度(DoM),即递归层级;2)在给定推理深度下进行正确推断。本文梳理了人工智能领域中多个方向的工作,包括大模型评测、心智理论增强模块、心智理论探针以及形式化心智理论模型。我们指出,当前研究主要集中于第二步,且多以静态逻辑问题为框架。为此,本文呼吁采用更贴近认知科学实验的动态环境来改进对心智理论能力的评估。

原文摘要 · Abstract (English)

Theory of Mind (ToM) capabilities in LLMs have recently become a central object of investigation. Cognitive science distinguishes between two steps required for ToM tasks: 1) determine whether to invoke ToM, which includes the appropriate Depth of Mentalizing (DoM), or level of recursion required to complete a task; and 2) applying the correct inference given the DoM. In this position paper, we first identify several lines of work in different communities in AI, including LLM benchmarking, ToM add-ons, ToM probing, and formal models for ToM. We argue that recent work in AI tends to focus exclusively on the second step which are typically framed as static logic problems. We conclude with suggestions for improved evaluation of ToM capabilities inspired by dynamic environments used in cognitive tasks.

心智理论大模型评估认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。