构建统一基准提升大模型共情理解能力
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
- 融合多轮互动与动态故事场景设计
- GPT-4o等模型在情绪信念任务中准确率超80%
- 适合研究大模型社会认知与人机共情的学者
理论心智(ToM)指理解自我与他人心理状态的能力,仍是大语言模型(LLMs)的挑战领域,其对人类心理状态的预测常不准确。本文提出UniToMBench,一个整合SimToM与TOMBENCH优势的统一基准,通过多轮互动任务设计与演进式故事场景,系统性地提升和评估LLMs的ToM能力。基于超过1,000条人工编写的情境数据集,UniToMBench结合视角转换技术与多样化评估指标,更有效激发模型的社会认知。评估显示,尽管GPT-4o与GPT-4o Mini在情感与信念相关任务中表现稳定,准确率普遍高于80%,但在知识型任务上仍存在显著性能波动。该结果凸显当前大模型在ToM任务中的优势与局限,证明UniToMBench作为未来发展的综合性评估工具具有重要价值。代码已开源:https://github.com/Shamant/unifiedtombenchmark。
原文摘要 · Abstract (English)
Theory of Mind (ToM), the ability to understand the mental states of oneself and others, remains a challenging area for large language models (LLMs), which often fail to predict human mental states accurately. In this paper, we introduce UniToMBench, a unified benchmark that integrates the strengths of SimToM and TOMBENCH to systematically improve and assess ToM capabilities in LLMs by integrating multi-interaction task designs and evolving story scenarios. Supported by a custom dataset of over 1,000 hand-written scenarios, UniToMBench combines perspective-taking techniques with diverse evaluation metrics to better stimulate social cognition in LLMs. Through evaluation, we observe that while models like GPT-4o and GPT-4o Mini show consistently high accuracy in tasks involving emotional and belief-related scenarios, with results usually above 80%, there is significant variability in their performance across knowledge-based tasks. These results highlight both the strengths and limitations of current LLMs in ToM-related tasks, underscoring the value of UniToMBench as a comprehensive tool for future development. Our code is publicly available here: https://github.com/Shamant/unifiedtombenchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。