构建社交推理新基准,评估大模型在复杂社会场景中的理解与判断能力。
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
- 设计六类任务,覆盖社交游戏、日常互动与数字社区三类真实场景。
- 模型在动态交互和信息不确定性下表现显著下降,链式思维强的模型更优。
- 通过针对性微调可大幅提升模型在复杂社交场景中的推理能力。
大型语言模型(LLMs)在在线社区治理、媒体内容分析及社交推理游戏等社会性任务中应用日益广泛。其成功依赖于社会推理能力——即解读社会情境、推断他人心理状态以及评估信息真实性。然而,当前缺乏系统评估框架来全面衡量LLM的社会推理能力。现有工作常简化现实场景,任务过于基础,难以挑战先进模型。为此,我们提出SocialMaze,一个专为评估社会推理能力而设计的新基准。该基准系统性地融合深度推理、动态交互与信息不确定性三大核心挑战,涵盖六种多样化任务,分布在社交推理游戏、日常生活互动和数字社区平台三个关键场景中。采用自动化与人工双重验证确保数据质量。评估发现:模型在处理动态交互与时间演化信息方面差异显著;具备强链式思维能力的模型在深层推理任务中表现更优;模型在不确定性环境下推理能力明显退化。此外,针对精选推理样本进行微调可显著提升模型在复杂社交场景中的表现。数据集已公开:https://huggingface.co/datasets/MBZUAI/SocialMaze。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning ability - the capacity to interpret social contexts, infer others' mental states, and assess the truthfulness of presented information. However, there is currently no systematic evaluation framework that comprehensively assesses the social reasoning capabilities of LLMs. Existing efforts often oversimplify real-world scenarios and consist of tasks that are too basic to challenge advanced models. To address this gap, we introduce SocialMaze, a new benchmark specifically designed to evaluate social reasoning. SocialMaze systematically incorporates three core challenges: deep reasoning, dynamic interaction, and information uncertainty. It provides six diverse tasks across three key settings: social reasoning games, daily-life interactions, and digital community platforms. Both automated and human validation are used to ensure data quality. Our evaluation reveals several key insights: models vary substantially in their ability to handle dynamic interactions and integrate temporally evolving information; models with strong chain-of-thought reasoning perform better on tasks requiring deeper inference beyond surface-level cues; and model reasoning degrades significantly under uncertainty. Furthermore, we show that targeted fine-tuning on curated reasoning examples can greatly improve model performance in complex social scenarios. The dataset is publicly available at: https://huggingface.co/datasets/MBZUAI/SocialMaze
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。