arXiv:2508.16072cs.AIcs.CL2025-08EMNLP被引 3

用推理游戏评测大模型能否模仿人类个性化思考方式

InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles

  • 构建包含策略记录与反思的评估框架,模拟人类推理风格
  • 11个主流大模型在《誓约》游戏中普遍依赖词汇线索,难随局势调整
  • 强化推理能力的模型已初现风格敏感性,适合研究人机认知对齐

大模型在以人为本的推理任务中表现优异,但现有评估多关注意图推断或欺骗检测,忽视了影响人际互动的个体化推理风格。社会推理游戏(SDGs)为评测此类风格提供了天然场景,同一条件下玩家可采用不同但合理的推理策略。为此,我们提出InMind——一个基于认知理论的评估框架,用于检验大模型在SDGs中捕捉并应用个性化推理风格的能力。InMind通过观察者与参与者模式收集回合级策略轨迹与赛后反思,支持四项认知驱动任务,联合评估静态匹配与动态适应能力。以《誓约》为例,评估11个前沿大模型发现:通用型模型(包括GPT-4o)频繁依赖词汇线索,难以关联时间序列游戏进程或动态调整策略;而强化推理的模型如DeepSeek-R1已展现出早期风格敏感迹象。结果揭示当前大模型在个性化、自适应推理上的关键局限,并推动向认知对齐的人机交互发展。

原文摘要 · Abstract (English)

LLMs have shown strong performance on human-centric reasoning tasks. While previous evaluations have explored whether LLMs can infer intentions or detect deception, they often overlook the individualized reasoning styles that influence how people interpret and act in social contexts. Social deduction games (SDGs) provide a natural testbed for evaluating individualized reasoning styles, where different players may adopt diverse but contextually valid reasoning strategies under identical conditions. To address this, we introduce InMind, a cognitively grounded evaluation framework designed to assess whether LLMs can capture and apply personalized reasoning styles in SDGs. InMind enhances structured gameplay data with round-level strategy traces and post-game reflections, collected under both Observer and Participant modes. It supports four cognitively motivated tasks that jointly evaluate both static alignment and dynamic adaptation. As a case study, we apply InMind to the game Avalon, evaluating 11 state-of-the-art LLMs. General-purpose LLMs, even GPT-4o frequently rely on lexical cues, struggling to anchor reflections in temporal gameplay or adapt to evolving strategies. In contrast, reasoning-enhanced LLMs like DeepSeek-R1 exhibit early signs of style-sensitive reasoning. These findings reveal key limitations in current LLMs' capacity for individualized, adaptive reasoning, and position InMind as a step toward cognitively aligned human-AI interaction.

大模型评估推理风格人机交互社会推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。