arXiv:2510.11254cs.CL2025-10Conference of the …被引 13

测试大模型的偏见,发现人类心理测验不靠谱。

Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality

  • 用人类心理测试评估17个大模型的性别、种族和道德偏见。
  • 测试结果与实际行为不匹配,甚至出现负相关。
  • 提醒直接套用人类测验于模型会误导结论。

心理测量测试被越来越多用于评估大语言模型(LLMs)的心理特质。然而,这些原本为人类设计的测试在应用于LLMs时是否有效仍不明确。本研究系统评估了17个LLMs在性别歧视、种族歧视和道德三个构念上的心理测量测试的信度与效度。结果显示,在多种题项和提示变体下具有中等信度。效度通过收敛效度(基于理论的跨测试相关性)和生态效度(测试得分与真实下游任务行为的一致性)进行评估。关键发现是:测试分数与下游任务中的模型行为不一致,部分情况下甚至呈负相关,表明其生态效度较低。结果强调,在解释测试分数前必须对心理测量工具在LLMs上的适用性进行系统评估,且人类心理测验不能不经调整直接用于大模型。

原文摘要 · Abstract (English)

Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs). However, it remains unclear whether these tests -- originally developed for humans -- yield meaningful results when applied to LLMs. In this study, we systematically evaluate the reliability and validity of human psychometric tests on 17 LLMs for three constructs: sexism, racism, and morality. We find moderate reliability across multiple item and prompt variations. Validity is evaluated through both convergent (i.e., testing theory-based inter-test correlations) and ecological approaches (i.e., testing the alignment between tests scores and behavior in real-world downstream tasks). Crucially, we find that psychometric test scores do not align, and in some cases even negatively correlate with, model behavior in downstream tasks, indicating low ecological validity. Our results highlight that systematic evaluations of psychometric tests on LLMs are essential before interpreting their scores. Our findings also suggest that psychometric tests designed for humans cannot be applied directly to LLMs without adaptation.

大模型评测偏见检测心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。