arXiv:2603.13126q-bio.NCcs.AI2026-03中稿 · August 2025被引 1

打造心理认知测评平台,系统评估大模型能力

Developing the PsyCogMetrics AI Lab to Evaluate Large Language Models and Advance Cognitive Science -- A Three-Cycle Action Design Science Study

  • 构建云上平台,融合心理测量与认知科学方法
  • 通过三轮设计验证,形成可复用的评测框架
  • 适合跨学科研究者,尤其关注AI与认知交叉领域

本研究提出并开发了PsyCogMetrics AI实验室(psycogmetrics.ai),一个集成化的云平台,用于将心理测量学和认知科学方法应用于大型语言模型(LLM)的评估。研究采用三周期行动设计科学范式:在相关性周期中识别现有评估方法的局限与利益相关者未满足需求;在严谨性周期中基于波普尔可证伪性、经典测验理论和认知负荷理论等核心理论,推导出演绎式设计目标;在设计周期中通过嵌套的构建-干预-评估循环实现这些目标。研究贡献了一个新颖的IT工具,以及一套经过验证的LLM评估设计,为人工智能、心理学、认知科学及社会科学交叉领域的研究提供支持。

原文摘要 · Abstract (English)

This study presents the development of the PsyCogMetrics AI Lab (psycogmetrics.ai), an integrated, cloud-based platform that operationalizes psychometric and cognitive-science methodologies for Large Language Model (LLM) evaluation. Framed as a three-cycle Action Design Science study, the Relevance Cycle identifies key limitations in current evaluation methods and unfulfilled stakeholder needs. The Rigor Cycle draws on kernel theories such as Popperian falsifiability, Classical Test Theory, and Cognitive Load Theory to derive deductive design objectives. The Design Cycle operationalizes these objectives through nested Build-Intervene-Evaluate loops. The study contributes a novel IT artifact, a validated design for LLM evaluation, benefiting research at the intersection of AI, psychology, cognitive science, and the social and behavioral sciences.

大模型评估认知科学心理测量AI平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。