用心理测量学评估大模型的心理推理能力,验证其可靠性。
AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities
- 引入心理测量方法,评估大模型的思维逻辑一致性。
- GPT-4和LLaMA-3在各项有效性指标上均优于前代模型。
- 为大模型可解释性提供新思路,适合关注AI可信性的研究者。
大规模语言模型(LLMs)参数量庞大、结构复杂,其表现类似人脑,却难以解释与评估。人工智能心理测量学(AI Psychometrics)是新兴领域,通过心理测量方法评估和解读人工智能的心理特征与认知过程。本文基于技术接受模型(TAM),对GPT-3.5、GPT-4、LLaMA-2和LLaMA-3四款主流大模型的心理推理能力与心理测量效度进行研究,考察了收敛效度、区分效度、预测效度与外部效度。结果表明,所有模型的响应均满足各项效度标准;性能更高的模型如GPT-4和LLaMA-3,在整体心理测量效度上显著优于其前代模型。研究证实了将人工智能心理测量学应用于大模型评估的可行性与有效性。
原文摘要 · Abstract (English)
The immense number of parameters and deep neural networks make large language models (LLMs) rival the complexity of human brains, which also makes them opaque ``black box'' systems that are challenging to evaluate and interpret. AI Psychometrics is an emerging field that aims to tackle these challenges by applying psychometric methodologies to evaluate and interpret the psychological traits and processes of artificial intelligence (AI) systems. This paper investigates the application of AI Psychometrics to evaluate the psychological reasoning and overall psychometric validity of four prominent LLMs: GPT-3.5, GPT-4, LLaMA-2, and LLaMA-3. Using the Technology Acceptance Model (TAM), we examined convergent, discriminant, predictive, and external validity across these models. Our findings reveal that the responses from all these models generally met all validity criteria. Moreover, higher-performing models like GPT-4 and LLaMA-3 consistently demonstrated superior psychometric validity compared to their predecessors, GPT-3.5 and LLaMA-2. These results help to establish the validity of applying AI Psychometrics to evaluate and interpret large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。