构建新基准,让模型学会用对话影响他人心理状态。
Understanding Mental States to Guide Social Influence in Multi-Person Group Dialogue
- 设计五幕情境,让模型扮演角色通过对话改变他人心理轨迹。
- 十款主流大模型平均表现比人类低54.2%,暴露长期心理建模短板。
- 适合研究社会智能、对话策略与心智理论的学者参考。
现有动态心智理论(ToM)评估多让语言模型被动观察:模型读取一系列连贯情境,报告人物信念、情绪、意图和行为变化。但在真实社交中,心智理论还用于主动行动——说话者规划言辞以引导他人心理状态朝目标演变。我们提出SocialMindChange基准,从追踪心理转向干预心理。每例包含四位角色和五个连续场景,模型扮演其中一人,在五幕中生成对话达成目标,同时保持所有参与者心理状态的一致性演化。基准还涵盖高阶心理状态。采用四步结构化框架,构建1200个社交情境,覆盖6000个场景和超过9万个问题,均经真实性与质量验证。对十款先进大模型的评估显示,其平均性能比人类低54.2%。这一差距表明当前大模型在长序列联动交互中仍难以维持并操控心理状态表征。
原文摘要 · Abstract (English)
Existing dynamic Theory of Mind (ToM) benchmarks mostly place language models in a passive role: the model reads a sequence of connected scenarios and reports what people believe, feel, intend, and do as these states change. In real social interaction, ToM is also used for action: a speaker plans what to say in order to shift another person's mental-state trajectory toward a goal. We introduce SocialMindChange, a benchmark that moves from tracking minds to changing minds in social interaction. Each instance defines a social context with 4 characters and five connected scenes. The model plays one character and generates dialogue across the five scenes to reach the target while remaining consistent with the evolving states of all participants. SocialMindChange also includes selected higher-order states. Using a structured four-step framework, we construct 1,200 social contexts, covering 6000 scenarios and over 90,000 questions, each validated for realism and quality. Evaluations on ten state-of-the-art LLMs show that their average performance is 54.2% below human performance. This gap suggests that current LLMs still struggle to maintain and change mental-state representations across long, linked interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。