arXiv:2410.08545cs.CLcs.AI2024-10被引 3

用文本挖掘检测大模型人格,更准且不受选项顺序干扰

Humanity in AI: Detecting the Personality of Large Language Models

  • 结合问卷与文本挖掘,避开选项顺序影响和幻觉问题
  • 发现ChatGPT、ChatGLM等具‘尽责性’人格特质,与人类接近
  • 预训练数据决定模型人格,指令微调能暴露隐藏性格特征

问卷是检测大语言模型(LLMs)人格的常用方法,但易受幻觉和选项顺序敏感性影响。为此,本文提出结合文本挖掘与问卷的方法:文本挖掘可提取模型回复中的心理特征,不受选项顺序干扰;且无需依赖特定答案,降低幻觉影响。通过归一化两方法得分并计算均方根误差,实验验证了该方法的有效性。在BERT、GPT等预训练模型(PLMs)及ChatGPT、ChatGLM等对话模型(ChatLLMs)上测试发现,LLMs具有明确人格特质,如ChatGPT和ChatGLM表现出‘尽责性’。进一步发现,模型人格源自其预训练数据;指令数据能增强人格生成,揭示隐藏性格。与人类平均人格分数对比,FLAN-T5(PLMs)和ChatGPT(ChatLLMs)得分差异分别为0.34和0.22,最接近人类。

原文摘要 · Abstract (English)

Questionnaires are a common method for detecting the personality of Large Language Models (LLMs). However, their reliability is often compromised by two main issues: hallucinations (where LLMs produce inaccurate or irrelevant responses) and the sensitivity of responses to the order of the presented options. To address these issues, we propose combining text mining with questionnaires method. Text mining can extract psychological features from the LLMs' responses without being affected by the order of options. Furthermore, because this method does not rely on specific answers, it reduces the influence of hallucinations. By normalizing the scores from both methods and calculating the root mean square error, our experiment results confirm the effectiveness of this approach. To further investigate the origins of personality traits in LLMs, we conduct experiments on both pre-trained language models (PLMs), such as BERT and GPT, as well as conversational models (ChatLLMs), such as ChatGPT. The results show that LLMs do contain certain personalities, for example, ChatGPT and ChatGLM exhibit the personality traits of 'Conscientiousness'. Additionally, we find that the personalities of LLMs are derived from their pre-trained data. The instruction data used to train ChatLLMs can enhance the generation of data containing personalities and expose their hidden personality. We compare the results with the human average personality score, and we find that the personality of FLAN-T5 in PLMs and ChatGPT in ChatLLMs is more similar to that of a human, with score differences of 0.34 and 0.22, respectively.

人格检测大模型文本挖掘心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。