arXiv:2604.16009cs.AI2026-04

测试大模型在社交压力下如何改变观点,揭示其自我反思能力差异。

MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models

论文配图:MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
图 1 · 摘自论文原文
  • 设计新评测框架,对比模型私密自审与受外部影响的修正行为。
  • 35个模型中多数在评估映射维度表现最差,显示对社会共识敏感度不一。
  • 适合研究大模型认知偏差、社会影响机制或人机信任的学者使用。

现有大模型评测多关注最终答案质量,却难以揭示模型在意见冲突或矛盾证据下如何修正信念。本文提出MEDLEY-BENCH,一个开放基准,通过统一基线比较结构化私密自审与分析师引导的社会修正行为。在130个实例上评估了来自12个模型家族的35个模型,采用梅德利元认知评分(MMS)及四个对齐分类体系、基于评分标准的复合指标。MMS点估计在同家族规模或生成版本间未呈现稳定排序。在预设内参照程序下,30/35模型在评估映射复合指标中得分最低;自我调节最低出现在4个模型,控制力最低1个。该评分依赖性模式可能反映模型行为、评判严格度、数据管道效应或其组合,非绝对评估缺陷证据。对11个精选模型的探索性对抗分析显示,对操纵共识标签的敏感度从接近零到显著响应不等。24组配对案例的人类评分研究显示,平均复合分差为0.727(95%置信区间:0.500–0.942)。12个共享情景下的24项交叉评分中,二次加权评委一致性kappa=0.389,人类-模型响应分布趋同度rho=0.637。本报告为MEDLEY-BENCH v1.0审计版概念验证发布。独立版本v1.5将重新执行完整协议,修正社交摘要呈现方式,并增强评分可复现性、实验控制与不确定性分析。MEDLEY-BENCH补充传统准确率评估,刻画模型在模糊与社会分歧情境下的有意识信念修正过程。

原文摘要 · Abstract (English)

Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence. We introduce MEDLEY-BENCH, an open benchmark comparing structured private self-review and analyst-conditioned social revision from a common solo baseline. We evaluated 35 models from 12 families on 130 instances using the Medley Metacognition Score (MMS) and four taxonomy-aligned, rubric-derived composites. MMS point estimates were not consistently ordered in the available within-family size or generation comparisons. Under the prespecified ipsative procedure, the Evaluation-mapped composite had the lowest relative rubric score in 30 of 35 models; Self-regulation was lowest in four models and Control in one. This rubric- and centering-dependent pattern may reflect model behavior, judge severity, data-pipeline effects, or their combination; it is not evidence of an absolute Evaluation deficit. In an exploratory adversarial analysis of 11 purposively selected models, sensitivity to manipulated consensus labels ranged from near zero to larger response shifts. A preliminary human rubric-application study of 24 paired vignettes found a mean composite difference of 0.727 (95% CI: 0.500-0.942). Across 24 cross-rated response items from 12 shared vignettes, quadratic-weighted inter-reviewer agreement was kappa = 0.389, and response-profile human-LLM convergence was rho = 0.637. This preprint reports MEDLEY-BENCH v1.0, the audited proof-of-concept release. A separately versioned v1.5 will rerun the full protocol with corrected social-summary rendering and stronger scoring reproducibility, experimental control, and uncertainty analysis. MEDLEY-BENCH complements accuracy-based evaluation by characterizing prompted belief revision under ambiguity and social disagreement.

元认知社会影响模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。