为大模型心理研究建立双重验证框架,避免误把幻象当现象。
From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
- 提出双有效性框架,融合心理测量与因果推断
- 不同研究目标需匹配相应证据强度,从文本分类到认知建模递进
- 强调构建计算版心理概念,而非直接套用人类量表
大语言模型(LLMs)正作为工具和研究对象进入心理学领域。然而,许多研究在未验证输出可靠性与可解释性的前提下,直接应用人类心理测量工具于模型,存在测量幻象风险——将统计规律误认为真实心理现象。本文主张,可靠的AI心理学研究需融合两种方法传统:一是心理测量学中对评分含义的验证,二是因果推断中对结果推论的严格标准。为此提出双有效性框架,其证据要求随科学目标提升而增强:从工具使用、行为表征、人类模拟到认知建模。文本分类仅需准确性和可靠性;声称模型模拟焦虑或揭示认知机制,则需额外证据,包括构念效度与实验控制。研究进展依赖于开发心理构念的计算对应物,而非假设人类量表可自动适用于语言模型。
原文摘要 · Abstract (English)
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms--statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。