构建对话推荐系统中评估大模型心智理论的新基准
RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender Systems
- 设计双维度评测框架:认知推断与行为预测
- 现有大模型在动态对话中难以持续保持策略性推理
- 适合研究对话智能与社会认知的学者使用
大型语言模型正通过其出色的指令理解、推理与人机交互能力,革新对话推荐系统。有效推荐对话的核心在于推断和理解用户心理状态(如需求、意图和信念),这一能力被称为心智理论(ToM)。尽管对大模型ToM评估的兴趣日益增长,现有基准多依赖受萨莉-安妮测试启发的合成叙事,侧重物理感知,未能捕捉真实对话场景中的心理状态推断复杂性。此外,现有评测常忽略人类ToM的关键组成部分——行为预测,即利用推断出的心理状态指导策略决策并选择未来对话行为的能力。为更贴近人类社会推理,我们提出RecToM,一个面向推荐对话中大模型心智理论能力的新基准。RecToM聚焦两个互补维度:认知推断(理解已传达信息并推断潜在心理状态)与行为预测(评估模型能否基于推断出的心理状态,预测、选择并评估适当的对话策略)。对先进大模型的广泛实验表明,RecToM构成重大挑战:模型虽部分具备识别心理状态的能力,但在动态推荐对话中难以维持连贯且具有战略性的心智推理,尤其在追踪不断变化的意图及使对话策略与推断心理状态一致方面表现不佳。
原文摘要 · Abstract (English)
Large Language models are revolutionizing the conversational recommender systems through their impressive capabilities in instruction comprehension, reasoning, and human interaction. A core factor underlying effective recommendation dialogue is the ability to infer and reason about users' mental states (such as desire, intention, and belief), a cognitive capacity commonly referred to as Theory of Mind. Despite growing interest in evaluating ToM in LLMs, current benchmarks predominantly rely on synthetic narratives inspired by Sally-Anne test, which emphasize physical perception and fail to capture the complexity of mental state inference in realistic conversational settings. Moreover, existing benchmarks often overlook a critical component of human ToM: behavioral prediction, the ability to use inferred mental states to guide strategic decision-making and select appropriate conversational actions for future interactions. To better align LLM-based ToM evaluation with human-like social reasoning, we propose RecToM, a novel benchmark for evaluating ToM abilities in recommendation dialogues. RecToM focuses on two complementary dimensions: Cognitive Inference and Behavioral Prediction. The former focus on understanding what has been communicated by inferring the underlying mental states. The latter emphasizes what should be done next, evaluating whether LLMs can leverage these inferred mental states to predict, select, and assess appropriate dialogue strategies. Extensive experiments on state-of-the-art LLMs demonstrate that RecToM poses a significant challenge. While the models exhibit partial competence in recognizing mental states, they struggle to maintain coherent, strategic ToM reasoning throughout dynamic recommendation dialogues, particularly in tracking evolving intentions and aligning conversational strategies with inferred mental states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。