arXiv:2503.12080cs.HCcs.AI2025-03被引 4

用大模型评估人格测试题目的内容效度,发现人机各有优势。

Comparing Human Expertise and Large Language Models Embeddings in Content Validity Assessment of Personality Tests

  • 对比人类专家与大模型对心理测试题的语义匹配判断。
  • 大模型在简洁题项上表现更优,人类在丰富行为描述题上更强。
  • 混合系统结合人机优势,可提升测评开发效率与客观性。

本文探讨大型语言模型(LLMs)在心理测量工具内容效度评估中的应用,聚焦于大五问卷(BFQ)和大五量表(BFI)。内容效度是测验构建的核心,确保心理测量工具充分覆盖目标构念。研究采用研究生心理学专家使用内容效度比(CVR)评分作为人类基准,同时利用多语言及微调的大模型分析题项嵌入向量,预测其与构念的映射关系。结果表明,人类验证者在行为描述丰富的BFQ题项上表现更佳,而大模型在语言简洁的BFI题项上更具优势。训练策略显著影响模型性能,针对词汇关系优化的模型优于通用大模型。研究凸显了融合人类专业知识与人工智能精度的混合验证系统的互补潜力,为心理测评的可扩展、客观且稳健的开发方法铺平道路。

原文摘要 · Abstract (English)

In this article we explore the application of Large Language Models (LLMs) in assessing the content validity of psychometric instruments, focusing on the Big Five Questionnaire (BFQ) and Big Five Inventory (BFI). Content validity, a cornerstone of test construction, ensures that psychological measures adequately cover their intended constructs. Using both human expert evaluations and advanced LLMs, we compared the accuracy of semantic item-construct alignment. Graduate psychology students employed the Content Validity Ratio (CVR) to rate test items, forming the human baseline. In parallel, state-of-the-art LLMs, including multilingual and fine-tuned models, analyzed item embeddings to predict construct mappings. The results reveal distinct strengths and limitations of human and AI approaches. Human validators excelled in aligning the behaviorally rich BFQ items, while LLMs performed better with the linguistically concise BFI items. Training strategies significantly influenced LLM performance, with models tailored for lexical relationships outperforming general-purpose LLMs. Here we highlights the complementary potential of hybrid validation systems that integrate human expertise and AI precision. The findings underscore the transformative role of LLMs in psychological assessment, paving the way for scalable, objective, and robust test development methodologies.

心理测评大模型内容效度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。