测试大模型能否从心理状态推断对话走向,发现人类远超AI。
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

- 用孤立心理状态预测对话轨迹,检验模型社会推理能力。
- 大模型能猜心理状态但难推对话发展,人类准确率100%。
- 适合研究具身智能、人机交互与心智理论的学者参考。
我们提出DialToM,一个基于真实人类对话构建的理论心智(ToM)基准,采用多选题评估框架。针对近期研究揭示的显式心理状态推断与实际应用中心智理论之间的差距,我们设计了更严格的「状态驱动诊断探针」,要求模型仅凭孤立的心理状态描述预测连贯的对话轨迹,不依赖上下文。评估显示,大模型在心理状态推断(字面意义的ToM)上表现良好,但在利用这些状态进行社交预测(功能性ToM)时严重不足。值得注意的是,领域专家在该任务上达到100%准确率,验证了任务的有效性,并揭示了显著的人机能力鸿沟。进一步实验表明,通过教师-学生推理注入机制,Gemini 3 Pro展现出强大的无上下文预测能力,且可迁移至较弱模型。DialToM的数据集、评估代码已公开于https://github.com/Stealth-py/DialToM。
原文摘要 · Abstract (English)
We introduce DialToM, an annotated Theory of Mind (ToM) benchmark built from naturalistic human-human dialogues using a multiple-choice evaluation framework. Concurrent with recent work showing a gap between explicit mental-state inference and applied ToM in synthetic settings~\cite{gu2024simpletom}, we establish a stricter \emph{State-Driven Diagnostic Probe} in which models must forecast state-consistent dialogue trajectories solely from isolated mental-state profiles without dialogue context. Our evaluation reveals a systematic reasoning asymmetry -- LLMs excel at inferring mental states (Literal ToM) but struggle to leverage them for social forecasting (Functional ToM). Crucially, a domain expert achieves 100\% accuracy on this task, proving its validity and establishing a stark human-AI capability gap. Further, a teacher-student reasoning injection probe shows that Gemini 3 Pro -- which establishes the leading baseline -- possesses robust Functional ToM capabilities for context-free forecasting that are transferable to weaker models. DialToM, its evaluation code, and dataset are publicly available at https://github.com/Stealth-py/DialToM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。