用心理量表评估大模型潜在人格特质,发现其与人类心理构念高度一致。
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
- 将心理量表转化为自然语言推理提示,实现对任意模型的心理测评
- 在88个模型中发现焦虑、抑郁等与人类相似的心理构念
- 为可解释、可控、可信的模型开发提供心理学工具支持
近期研究表明大型语言模型中存在类人个性特征,暗示其已知及未知偏见可能符合人类潜在心理构念。尽管大型对话模型可通过欺骗方式回答心理量表,但数千个用于其他任务的简化Transformer模型因缺乏合适心理测量方法而难以评估。本文提出将标准心理量表重构为自然语言推理提示,并发布代码库以支持任意模型的心理测量。通过对88个公开模型的样本分析,证实了与人类心理健康相关构念(包括焦虑、抑郁和整体意义感)的存在,这些构念符合人类心理学标准理论,表现出类似的相关性与缓解策略。利用心理工具解读和修正语言模型性能,有助于推动更可解释、可控且可信模型的发展。
原文摘要 · Abstract (English)
Human-like personality traits have recently been discovered in large language models, raising the hypothesis that their (known and as yet undiscovered) biases conform with human latent psychological constructs. While large conversational models may be tricked into answering psychometric questionnaires, the latent psychological constructs of thousands of simpler transformers, trained for other tasks, cannot be assessed because appropriate psychometric methods are currently lacking. Here, we show how standard psychological questionnaires can be reformulated into natural language inference prompts, and we provide a code library to support the psychometric assessment of arbitrary models. We demonstrate, using a sample of 88 publicly available models, the existence of human-like mental health-related constructs (including anxiety, depression, and Sense of Coherence) which conform with standard theories in human psychology and show similar correlations and mitigation strategies. The ability to interpret and rectify the performance of language models by using psychological tools can boost the development of more explainable, controllable, and trustworthy models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。