arXiv:2505.08245cs.CLcs.AI2025-05综述被引 49

用心理测量学方法评估大模型的人格与认知能力

Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement

  • 引入心理测量理论评估大模型的隐性心理特征
  • 构建涵盖人格、价值观等维度的多维评估框架
  • 适合关注人机对齐与可信AI的研究者参考

大语言模型的发展已超越传统评估方法,带来新挑战:如何衡量类人心理特质、突破静态任务基准、建立以人为中心的评估体系。这些挑战与心理测量学——量化人格、价值观、智力等无形心理特征的科学——高度契合。本文系统综述新兴的LLM心理测量学领域,整合心理测量工具、理论与原则,用于评估、理解与提升大模型。通过梳理文献,本文构建了基准化原则,拓展评估范围,优化方法,验证结果,并推动模型能力进步。多视角融合形成结构化框架,为跨学科研究者提供指引,助力发展符合人类水平的人工智能评估范式,促进以人为本的AI系统建设。相关资源库可在 https://github.com/valuebyte-ai/Awesome-LLM-Psychometrics 获取。

原文摘要 · Abstract (English)

The advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. This progress presents novel challenges, such as measuring human-like psychological constructs, moving beyond static and task-specific benchmarks, and establishing human-centered evaluation. These challenges intersect with psychometrics, the science of quantifying the intangible aspects of human psychology, such as personality, values, and intelligence. This review paper introduces and synthesizes the emerging interdisciplinary field of LLM Psychometrics, which leverages psychometric instruments, theories, and principles to evaluate, understand, and enhance LLMs. The reviewed literature systematically shapes benchmarking principles, broadens evaluation scopes, refines methodologies, validates results, and advances LLM capabilities. Diverse perspectives are integrated to provide a structured framework for researchers across disciplines, enabling a more comprehensive understanding of this nascent field. Ultimately, the review provides actionable insights for developing future evaluation paradigms that align with human-level AI and promote the advancement of human-centered AI systems for societal benefit. A curated repository of LLM psychometric resources is available at https://github.com/valuebyte-ai/Awesome-LLM-Psychometrics.

心理测量大模型评估人机对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。