评测大模型在说服对话中理解心理状态的能力
PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- 设计双任务框架评估模型对心理状态的追踪与应用
- 8个主流模型在动态心理状态追踪上表现不佳
- 适合研究社会智能与对话系统的学者参考
理解并预测自我与他人心理状态的能力,即心智理论(ToM),在有效社交场景中至关重要。尽管近期研究已评估大语言模型(LLMs)的ToM能力,但现有基准多聚焦于简化场景(如Sally-Anne类任务),忽视了真实社交互动的复杂性。为弥补这一差距,我们提出PersuasiveToM,一个专用于评估LLMs在说服对话中ToM能力的基准。该框架包含两项核心任务:ToM推理,测试对不断变化的欲望、信念和意图的追踪;ToM应用,评估基于推断心理状态来预测和评估说服策略的能力。在八个领先大模型上的实验表明,尽管模型在多项问题上表现良好,但在追踪心理状态动态变化及全面理解整段对话中的心理状态方面仍存在明显短板。PersuasiveToM旨在更关注复杂心理活动,实现对大模型ToM推理能力的有效评估。代码已开源。
原文摘要 · Abstract (English)
The ability to understand and predict the mental states of oneself and others, known as the Theory of Mind (ToM), is crucial for effective social scenarios. Although recent studies have evaluated ToM in Large Language Models (LLMs), existing benchmarks focus on simplified settings (e.g., Sally-Anne-style tasks) and overlook the complexity of real-world social interactions. To mitigate this gap, we propose PersuasiveToM, a benchmark designed to evaluate the ToM abilities of LLMs in persuasive dialogues. Our framework contains two core tasks: ToM Reasoning, which tests tracking of evolving desires, beliefs, and intentions; and ToM Application, which assesses the use of inferred mental states to predict and evaluate persuasion strategies. Experiments across eight leading LLMs reveal that while models excel on multiple questions, they struggle with the tasks that need tracking the dynamics and shifts of mental states and understanding the mental states in the whole dialogue comprehensively. Our aim with PersuasiveToM is to allow an effective evaluation of the ToM reasoning ability of LLMs with more focus on complex psychological activities. Our code is available at https://github.com/Yu-Fangxu/PersuasiveToM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。