arXiv:2510.11389cs.CL2025-10被引 12

用真人对战数据评估大模型在推理与策略上的社会智能表现

Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies

  • 以获胜方策略为标准,分语言立场与决策两阶段评估模型表现
  • 顶尖大模型仅半数得分超0.5,暴露其欺骗与反事实推理短板
  • 适合研究多智能体交互中语言与策略协同的学者参考

狼人杀等社交推理游戏融合语言、推理与策略,是检验自然语言与社会智能的理想场景。然而现有研究多依赖大模型自对弈,导致话语模板化、案例片面,且评估依赖生存时长或主观评分等粗粒度指标。为此,我们构建了一个高质量、人工验证的多模态狼人杀数据集,包含超100小时视频、3240万条语音文本及15种规则变体。基于该数据集,提出一种策略对齐评估框架:第一阶段为言语评估,通过多选题形式考察模型在五维社交能力中的立场适配性;第二阶段为决策评估,检验模型的投票选择与对手角色推断能力。该框架可细粒度衡量模型的语言与推理能力,并捕捉其策略连贯性。实验表明,当前主流大模型表现分化显著,约一半得分低于0.5,暴露出在欺骗与反事实推理方面的明显不足。期望本数据集推动语言、推理与策略在多智能体交互中的研究。

原文摘要 · Abstract (English)

Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies reduce the game to LLM-based self-play, yielding templated utterances and anecdotal cases that overlook the richness of social gameplay. Evaluation further relies on coarse metrics such as survival time or subjective scoring due to the lack of quality reference data. To address these gaps, we curate a high-quality, human-verified multimodal Werewolf dataset containing over 100 hours of video, 32.4M utterance tokens, and 15 rule variants. Based on this dataset, we propose a novel strategy-alignment evaluation that leverages the winning faction's strategies as ground truth in two stages: 1) Speech evaluation, formulated as multiple-choice-style tasks that assess whether the model can adopt appropriate stances across five dimensions of social ability; and 2) Decision evaluation, which assesses the model's voting choices and opponent-role inferences. This framework enables a fine-grained evaluation of models' linguistic and reasoning capabilities, while capturing their ability to generate strategically coherent gameplay. Our experiments show that state-of-the-art LLMs show diverse performance, with roughly half remain below 0.50, revealing clear gaps in deception and counterfactual reasoning. We hope our dataset further inspires research on language, reasoning, and strategy in multi-agent interaction.

社交推理大模型评估策略生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。