用程序生成复杂心理理论数据,揭露大模型真实能力短板
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
- 用A*搜索在自定义语言中生成复杂故事,构造多样化挑战场景
- GPT-4o和Llama-3.1-70B在新数据上准确率低至9%和0%
- 适合评估模型社会智能,尤其关注心理推理能力的科研人员
大语言模型是否具备心理理论能力?现有大量论文与评测基准多依赖简单模式的数据集,易导致评估盲区并高估模型表现。本文提出ExploreToM,首个支持大规模生成多样且具挑战性的心理理论数据的框架。该方法基于自定义领域特定语言,利用A*搜索生成复杂故事结构与新颖、多样且合理的场景,以全面测试大模型极限。评估显示,当前顶尖模型如Llama-3.1-70B和GPT-4o在生成数据上的准确率分别低至0%和9%,凸显现有评测的不足。由于生成数据是先前工作的概念超集,基于其微调可在经典ToMi基准(Le et al., 2019)上实现27个百分点的准确率提升。此外,该框架还揭示了模型缺乏可靠状态追踪或数据不平衡等关键缺陷,可能是其在基准上表现不佳的原因。
原文摘要 · Abstract (English)
Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited datasets with simple patterns that can potentially lead to problematic blind spots in evaluation and an overestimation of model capabilities. We introduce ExploreToM, the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation. Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs. Our evaluation reveals that state-of-the-art LLMs, such as Llama-3.1-70B and GPT-4o, show accuracies as low as 0% and 9% on ExploreToM-generated data, highlighting the need for more robust theory of mind evaluation. As our generations are a conceptual superset of prior work, fine-tuning on our data yields a 27-point accuracy improvement on the classic ToMi benchmark (Le et al., 2019). ExploreToM also enables uncovering underlying skills and factors missing for models to show theory of mind, such as unreliable state tracking or data imbalances, which may contribute to models' poor performance on benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。