arXiv:2605.06444cs.AI2026-05被引 2

评测大模型对社会概念的深度推理能力,发现顶尖模型已超越人类专家。

SCRuB: Social Concept Reasoning under Rubric-Based Evaluation

论文配图:SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
图 1 · 摘自论文原文
  • 基于五维批判思维量规,构建社会概念推理评估框架。
  • 在1170组对比中,模型胜出80.8%,整体偏好率达74.4%。
  • 适合关注社会认知、伦理推理与模型评估的研究者。

尽管大量研究聚焦大语言模型在数学或技术任务中的推理能力,但很少关注社会概念推理——即塑造社会规范、文化与制度的抽象理念。这一能力对作为社会代理的大模型至关重要,却缺乏系统评估方法。本文提出SCRuB(基于量规的社会概念推理评估框架),应对任务不确定性问题。目标是衡量模型在社会概念推理上是否达到人类专家的深度与批判严谨性。流程包含三阶段:从权威来源构建提示、专家与模型生成回应、使用五维批判思维量规进行对比评估。为提升泛化性,引入跨学科视角专家组,并经独立专家验证。发布SCRuBEval(4,711个评估提示)与SCRuBAnnotations(300条专家回应、150组专家对比判断,来自45位博士级学者)。结果显示,前沿模型在所有五维量规上均显著优于人类专家。在1,170组两两比较中,专家将模型回应排第一的比例达80.8%,整体更偏好模型回应74.4%。本研究首次实现社会概念推理的专家基准评估饱和:单轮问答式测试已触及模型与人类的上限。

原文摘要 · Abstract (English)

While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas shaping social norms, culture, and institutions. This understudied capability is essential for modern models acting as social agents, yet no systematic evaluation methodology targets it. We introduce SCRuB (Social Concept Reasoning under Rubric-Based Evaluation), a framework designed for this setting of task indeterminacy. Our goal is to measure the degree to which a model reasons about social concepts with the depth and critical rigor of a human expert. SCRuB proceeds in three phases: prompt construction from established sources, response generation by experts and models, and comparative evaluation using a five-dimensional critical thinking rubric. To enable generalization of the pipeline, we introduce a Panel of Disciplinary Perspectives ensemble validated against independent expert judges. We release SCRuBEval (n=4,711 evaluation prompts) and SCRuBAnnotations (300 expert-authored responses and 150 expert comparative judgments from 45 PhD-level scholars). Our results show that frontier models consistently outperform human experts across all five rubric dimensions. Across 1,170 pairwise comparisons, expert judges ranked a model response first in 80.8% of judgments and preferred model responses overall 74.4% of the time. Ultimately, this study provides the first expert-grounded demonstration of evaluation saturation for social concept reasoning: the single-turn exam-style format has reached its ceiling for models and humans alike.

社会推理评估框架大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。