评测大模型在200个情境下的动机推理能力,发现其仍远未达到人类水平。
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
- 构建包含200个情境、600个任务的MotivBench评测集
- 7大模型家族测试显示最先进模型仍难理解'爱与归属'类动机
- 揭示模型过度理性化倾向,适合研究AI人性化方向者参考
大型语言模型(LLMs)已被广泛用作社交模拟和人工智能伴侣等场景中的核心组件。然而,它们在多大程度上能复现人类动机仍缺乏深入探讨。现有基准受限于简单场景和缺乏角色身份,导致与现实情境存在信息不对称。为此,我们提出MotiveBench,包含200个丰富情境和600个推理任务,覆盖多重动机层次。基于该基准,我们在7个主流模型家族中开展大规模实验,比较各家族内不同规模与版本的表现。结果表明,即使最先进的模型在实现人类级动机推理方面仍明显不足。分析发现,模型在理解‘爱与归属’类动机时尤为困难,且表现出过度理性和理想主义倾向。这些发现为未来大模型人性化研究指明了方向。数据集、基准及代码已公开于https://aka.ms/motivebench。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely adopted as the core of agent frameworks in various scenarios, such as social simulations and AI companions. However, the extent to which they can replicate human-like motivations remains an underexplored question. Existing benchmarks are constrained by simplistic scenarios and the absence of character identities, resulting in an information asymmetry with real-world situations. To address this gap, we propose MotiveBench, which consists of 200 rich contextual scenarios and 600 reasoning tasks covering multiple levels of motivation. Using MotiveBench, we conduct extensive experiments on seven popular model families, comparing different scales and versions within each family. The results show that even the most advanced LLMs still fall short in achieving human-like motivational reasoning. Our analysis reveals key findings, including the difficulty LLMs face in reasoning about "love & belonging" motivations and their tendency toward excessive rationality and idealism. These insights highlight a promising direction for future research on the humanization of LLMs. The dataset, benchmark, and code are available at https://aka.ms/motivebench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。