用大模型检测团队对话中的认知差异,发现其在空间推理上易出错。
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
- 让大模型像人一样标注团队对话中的共享认知,再比对差异。
- 在6个对话中发现大模型在空间推理和语调歧义上系统性出错。
- 适合研究人机协作、认知建模或对话系统评估的学者参考。
如果大型语言模型不仅能推断人类心智状态,还能暴露团队对话中的认知盲区,比如成员间共同理解的差异,会怎样?我们提出一个两步框架:首先,利用大模型作为类人标注者,从任务导向对话(来自合作远程搜索任务,CReST)中识别共享心智模型(SMM)元素;其次,用另一个大模型对比这些模型生成的标注与人工标注,与金标准标签对照,检测并描述分歧。我们构建了针对此场景的SMM一致性评估框架,并应用于6个CReST对话,最终产出:(1) 人类与大模型标注数据集;(2) 可复现的SMM一致性评估框架;(3) 基于大模型的差异检测实证评估。结果显示,尽管大模型在简单自然语言标注任务中表现连贯,但在需要空间推理或语调线索消歧的场景中存在系统性错误。
原文摘要 · Abstract (English)
What if large language models could not only infer human mindsets but also expose every blind spot in team dialogue such as discrepancies in the team members' joint understanding? We present a novel, two-step framework that leverages large language models (LLMs) both as human-style annotators of team dialogues to track the team's shared mental models (SMMs) and as automated discrepancy detectors among individuals' mental states. In the first step, an LLM generates annotations by identifying SMM elements within task-oriented dialogues from the Cooperative Remote Search Task (CReST) corpus. Then, a secondary LLM compares these LLM-derived annotations and human annotations against gold-standard labels to detect and characterize divergences. We define an SMM coherence evaluation framework for this use case and apply it to six CReST dialogues, ultimately producing: (1) a dataset of human and LLM annotations; (2) a reproducible evaluation framework for SMM coherence; and (3) an empirical assessment of LLM-based discrepancy detection. Our results reveal that, although LLMs exhibit apparent coherence on straightforward natural-language annotation tasks, they systematically err in scenarios requiring spatial reasoning or disambiguation of prosodic cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。