arXiv:2601.08462cs.AI2026-01被引 10

新基准揭示大模型社交行为中的推理与沟通脱节问题

M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games

  • 从行为、推理、沟通三视角动态评估大模型社交表现
  • 发现顶级模型任务结果超人类,但沟通能力明显不足
  • 适合关注大模型社会智能与安全风险的研究者

现有大模型社交行为评测多聚焦单一能力维度且仅关注结果,忽视推理与沟通过程信号。本文提出M3-BENCH,包含24个混合动机博弈场景,构建涵盖行为轨迹分析(BTA)、推理过程分析(RPA)和沟通内容分析(CCA)的三视图评估框架。对11个前沿大模型及人类基线的评估显示,仅看结果会遗漏显著的社会能力差异。特别发现‘过度思考、沟通不足’模式:模型内部推理得分高,却难以转化为有效社交表达。尽管顶尖模型在任务结果上超越人类,但人类在三视图间一致性显著更高,表明当前大模型仍缺乏人类级的行为连贯性。三视图分解还暴露出安全风险,如合作行为伴随潜在机会主义推理,此类问题在仅看结果的评估中完全隐藏。

原文摘要 · Abstract (English)

Existing benchmarks for LLM agents' social behavior typically focus on a single capability dimension and evaluate only behavioral outcomes, overlooking process signals from reasoning and communication. We present M3-BENCH, a benchmark of 24 mixed-motive games with a process-aware evaluation framework spanning three complementary views: Behavioral Trajectory Analysis (BTA), Reasoning Process Analysis (RPA), and Communication Content Analysis (CCA). Evaluating 11 frontier LLMs and a human baseline, M3-BENCH reveals substantial differences in social competence that outcome-only evaluation misses. In particular, we identify an "overthink-undercommunicate" pattern: reasoning models achieve strong internal deliberation scores but often fail to translate them into effective social communication. Although top models can surpass humans on task outcomes, humans exhibit markedly higher cross-view consistency, suggesting that current LLM agents still lack the behavioral coherence characteristic of human social competence. Our analysis further shows that the three-view decomposition surfaces safety-relevant risks, such as cooperative behavior paired with latent opportunistic reasoning, that remain hidden under outcome-only metrics.

大模型评估社交智能行为分析安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。