用心理测评方法评估大模型自我报告的真实性,发现部分模型存在系统性失真。
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report

- 借鉴临床量表思路,设计六项有效性指标检测大模型自报告偏差。
- 4个模型被判定为无效水平,其置信度与回答正确性无关联(相关系数-0.20)。
- 适用于评估大模型可信度,尤其适合关注模型自我认知真实性的研究者。
临床人格评估在解读量表前会筛查反应有效性,而大模型评估尚未如此。本文将PAI和MMPI-3中的有效性量表框架应用于20个前沿大模型的524项元认知探测数据,构建六项有效性指标:L(错误中保持信心)、K(错误中下注)、F(回避共识项目)、Fp(回避正确答案)、RBS(监控反转)和TRIN(固定响应)。分级分类系统识别出4个模型为构念层面无效,2个为异常升高。有效模型表现出项目敏感的置信度(均相关系数r = .18,16项中有14项显著),无效模型则无此关联(均相关系数r = -.20,d = 2.17,p = .001)。链式思维训练导致两种相反的反应扭曲。两个潜在维度解释了94.6%的指标方差。配套论文提出可移植的筛查协议(Cacioli, 2026e)并验证其选择性预测能力(Cacioli, 2026f)。所有数据与代码见:https://github.com/synthiumjp/validity-scaling-llm
原文摘要 · Abstract (English)
Clinical personality assessment screens response validity before interpreting substantive scales. LLM evaluation does not. We apply the validity scaling framework from the PAI and MMPI-3 to metacognitive probe data from 20 frontier models across 524 items. Six validity indices are operationalised: L (maintaining confidence on errors), K (betting on errors), F (withdrawing consensus-endorsed items), Fp (withdrawing correct answers), RBS (inverted monitoring), and TRIN (fixed responding). A tiered classification system identifies four models as construct-level invalid and two as elevated. Valid-profile models produce item-sensitive confidence (mean r = .18, 14 of 16 significant). Invalid-profile models do not (mean r = -.20, d = 2.17, p = .001). Chain-of-thought training produces two opposite response distortions. Two latent dimensions account for 94.6% of index variance. Companion papers extract a portable screening protocol (Cacioli, 2026e) and validate it against selective prediction (Cacioli, 2026f). All data and code: https://github.com/synthiumjp/validity-scaling-llm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。