大模型评估中人类认知框架失效,导致智商评分与实际表现严重脱节。
The Catastrophic Paradox of Human Cognitive Frameworks in Large Language Model Evaluation: A Comprehensive Empirical Analysis of the CHC-LLM Incompatibility
- 用CHC智力理论评估9个前沿大模型,发现人类认知框架不适用。
- 模型智商评分85.0~121.4,但晶体知识任务准确率接近零。
- 适合关注AI评估方法论与认知偏见的研究者阅读。
本研究通过系统评估包括GPT-5、Claude Opus 4.1和Gemini 3 Pro Preview在内的九个前沿模型,基于卡特尔-霍恩-卡罗尔智力理论,揭示了人类心理测量框架与大模型评估之间的根本不兼容性。结果显示,模型智商得分在85.0至121.4之间时,其晶体知识任务的二元准确率趋近于零,评委评分与模型表现的相关系数r=0.175(p=0.001,n=1800)。该矛盾在晶体智力领域最为显著:所有模型均实现完美二元准确率,而评委评分仅在25%至62%之间,此现象在有效测量条件下不可能发生。通过项目反应理论建模、跨厂商评委验证及矛盾严重性指数分析,我们指出这种脱节源于将生物认知架构错误地应用于基于Transformer的系统。研究挑战了智能、测量与人工智能评估中的人类中心偏见,提出应发展适应机器自身特性的原生认知评估框架。
原文摘要 · Abstract (English)
This investigation presents an empirical analysis of the incompatibility between human psychometric frameworks and Large Language Model evaluation. Through systematic assessment of nine frontier models including GPT-5, Claude Opus 4.1, and Gemini 3 Pro Preview using the Cattell-Horn-Carroll theory of intelligence, we identify a paradox that challenges the foundations of cross-substrate cognitive evaluation. Our results show that models achieving above-average human IQ scores ranging from 85.0 to 121.4 simultaneously exhibit binary accuracy rates approaching zero on crystallized knowledge tasks, with an overall judge-binary correlation of r = 0.175 (p = 0.001, n = 1800). This disconnect appears most strongly in the crystallized intelligence domain, where every evaluated model achieved perfect binary accuracy while judge scores ranged from 25 to 62 percent, which cannot occur under valid measurement conditions. Using statistical analyses including Item Response Theory modeling, cross-vendor judge validation, and paradox severity indexing, we argue that this disconnect reflects a category error in applying biological cognitive architectures to transformer-based systems. The implications extend beyond methodology to challenge assumptions about intelligence, measurement, and anthropomorphic biases in AI evaluation. We propose a framework for developing native machine cognition assessments that recognize the non-human nature of artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。