构建心理测评基准,揭示大模型与人类在认知特性上的差异。
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
- 设计轻量角色扮演提示,提升大模型响应率至90.4%。
- 多语言测试显示,7种语言在43个维度上与英语偏差5%-20.2%。
- 首次系统揭示语言对大模型心理属性的影响,适合研究者参考。
具备数百亿参数的大语言模型(LLMs)通过学习海量互联网数据展现出类人智能,但大规模神经网络的不可解释性引发对其可靠性的担忧。现有研究借用人类心理学概念评估大模型的心理特质,却未考虑其与人类的根本差异,导致直接复用人类量表时出现高拒答率。此外,这些量表无法衡量不同语言下大模型心理属性的变化。本文提出AIPsychoBench,一个专为评估大模型心理特质设计的基准。该基准采用轻量级角色扮演提示,绕过模型对齐机制,将平均有效响应率从70.12%提升至90.40%,同时正负偏差分别仅为3.3%和2.1%,显著低于传统越狱提示造成的9.8%和6.9%偏差。在总计112个心理测量子类别中,七种语言相对于英文的得分偏差范围为5%至20.2%,覆盖43个子类别,提供了关于语言对大模型心理特质影响的首份全面证据。
原文摘要 · Abstract (English)
Large Language Models (LLMs) with hundreds of billions of parameters have exhibited human-like intelligence by learning from vast amounts of internet-scale data. However, the uninterpretability of large-scale neural networks raises concerns about the reliability of LLM. Studies have attempted to assess the psychometric properties of LLMs by borrowing concepts from human psychology to enhance their interpretability, but they fail to account for the fundamental differences between LLMs and humans. This results in high rejection rates when human scales are reused directly. Furthermore, these scales do not support the measurement of LLM psychological property variations in different languages. This paper introduces AIPsychoBench, a specialized benchmark tailored to assess the psychological properties of LLM. It uses a lightweight role-playing prompt to bypass LLM alignment, improving the average effective response rate from 70.12% to 90.40%. Meanwhile, the average biases are only 3.3% (positive) and 2.1% (negative), which are significantly lower than the biases of 9.8% and 6.9%, respectively, caused by traditional jailbreak prompts. Furthermore, among the total of 112 psychometric subcategories, the score deviations for seven languages compared to English ranged from 5% to 20.2% in 43 subcategories, providing the first comprehensive evidence of the linguistic impact on the psychometrics of LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。