构建可复现的语言模型心理测评框架,提升实验可靠性
R.U.Psycho? Robust Unified Psychometric Testing of Language Models
- 提出R.U.Psycho框架,简化语言模型心理测试的流程设计
- 在多种量表上验证框架有效性,结果与已有研究一致
- 无需复杂编程,适合社会科学研究者快速开展实验
生成式语言模型正被用于模拟人类心理测评,以评估其性格特征、对齐程度或作为社会科学实验中的虚拟参与者。然而,模型输出不稳定、提示设计敏感、参数设置多样及模型版本繁多,导致实验难以复现且结论泛化困难。本文提出R.U.Psycho框架,旨在实现语言模型心理测评的鲁棒性与可复现性,仅需有限编码能力即可运行。我们在多种心理测量量表上验证了该框架的有效性,结果支持文献中已有发现。R.U.Psycho已开源,可通过https://github.com/julianschelb/rupsycho获取。
原文摘要 · Abstract (English)
Generative language models are increasingly being subjected to psychometric questionnaires intended for human testing, in efforts to establish their traits, as benchmarks for alignment, or to simulate participants in social science experiments. While this growing body of work sheds light on the likeness of model responses to those of humans, concerns are warranted regarding the rigour and reproducibility with which these experiments may be conducted. Instabilities in model outputs, sensitivity to prompt design, parameter settings, and a large number of available model versions increase documentation requirements. Consequently, generalization of findings is often complex and reproducibility is far from guaranteed. In this paper, we present R.U.Psycho, a framework for designing and running robust and reproducible psychometric experiments on generative language models that requires limited coding expertise. We demonstrate the capability of our framework on a variety of psychometric questionnaires, which lend support to prior findings in the literature. R.U.Psycho is available as a Python package at https://github.com/julianschelb/rupsycho.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。