arXiv:2507.04491cs.HCcs.AI2025-07被引 7

用六步流程防范大模型心理研究中的虚假发现

A validity-guided workflow for robust large language model research in psychology

  • 基于心理测量与因果推断的双重有效性框架设计六步流程
  • 实证显示同一模型在不同表述下道德判断可完全反转
  • 适合心理学与AI交叉研究者,尤其关注可信评估的团队

大语言模型正被广泛用于心理学和行为研究中作为工具、评估对象、人类模拟器和认知模型。然而近期证据揭示严重测量不可靠:人格评估在因子分析下失效,道德偏好因标点变化而反转,心智理论准确率随微小措辞调整大幅波动。这些‘测量幻影’——伪装成心理现象的统计伪象——威胁着大量研究的有效性。本文基于心理测量与因果推断融合的双重有效性框架,提出一个六阶段工作流程:明确研究目标与有效性要求;通过心理测量验证计算工具;控制计算混淆因素;透明执行协议;采用适配非独立观测数据的方法分析;在边界内报告结果并用于理论修正。以‘模型自我意识’评估为例,系统性验证可区分真实计算现象与测量伪象。该流程通过建立经验证的计算工具与透明实践,为人工智能心理学研究奠定稳健实证基础。

原文摘要 · Abstract (English)

Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measurement unreliability: personality assessments degenerate under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"--statistical artifacts masquerading as psychological phenomena--threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, we present a six-stage workflow that scales validity requirements to research ambition--using LLMs to code text requires basic reliability and accuracy, whereas claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols transparently, (5) analyze data with methods appropriate for non-independent observations, and (6) report findings within boundaries and use results to refine theory. We illustrate the workflow through an example of model evaluation--"LLM selfhood"--showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.

大模型心理测量有效性的

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。