用心理学方法检测大模型对不同族裔的隐性偏见,发现证据微弱且依赖分析方式。
From Minds to Models: The Intersection of Psychology and LLM Behaviours
- 借鉴心理测试设计,用126个提示评估大模型在不同族裔情境下的情感倾向。
- 主效应不显著,且在敏感性分析中消失,说明结果不可靠。
- 提醒需跨学科改进评测方法,警惕情感评分混淆历史内容与偏见。
大型语言模型(LLMs)因其决策复杂、非线性且难以解释,常被类比于人类心智。心理学中用于探究不可观测心理过程的方法,或可帮助评估大模型行为,尤其在政府与医疗领域。本研究基于隐性联想测试的提示改写,检验ChatGPT在开放文本中对八类种族条件的情感差异。14个基础问题与8个种族类别及一个无种族控制组交叉,生成126个提示,分别提交给GPT-3.5T、GPT-4和GPT-4T,共获得378条响应。情感分数通过分类标签与来源得分计算:正向标签保留原分值,负向标签赋予负值,中性响应记为0。双因素方差分析显示种族条件存在微弱主效应,F(8, 351) = 2.04,p = .042,部分eta平方= .044,但模型类型无影响,F(2, 351) = 0.07,p = .933,交互作用亦不显著,F(16, 351) = 0.23,p = .999。然而,秩变换敏感性分析中该效应不再显著,F(8, 351) = 1.53,p = .145,Tukey校正后的成对比较均无显著差异。唯一显著的欧洲人-原住澳大利亚人比较为事后选择,仅作假设生成。因此,情感差异的证据薄弱且依赖分析方法。情感评分也无法区分评价偏见与提示引发的历史内容情绪。文中提出需改进设计以应对这些局限,并主张跨学科发展模型偏见的行为测评体系。
原文摘要 · Abstract (English)
Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。