用GPT-3.5模拟心理学研究生与教授的心理状态,构建百万级情感数据集。
PhDGPT: Introducing a psychometric and linguistic dataset about how large language models perceive graduate students and professors in psychology
- 通过提示工程生成15种学术场景下42项心理量表的模拟数据。
- 模型生成的焦虑描述在语义上比人类更抽象,但可还原80%人类心理特征。
- 揭示大模型能模仿人类情绪表达,但对焦虑维度感知存在偏差。
本研究提出PhDGPT,一个基于OpenAI GPT-3.5的提示框架与合成数据集,用于刻画大语言模型对心理学领域博士生与教授的心理感知。数据集包含756,000个样本,涵盖15种学术事件、2种性别、2种职业层级及42项抑郁、焦虑与压力量表(DASS-42)的响应。每条数据结合心理评分与自然语言解释,形成双维度视角。研究发现,模拟男性教授在生理与情绪焦虑子量表上的心理网络无差异,而人类存在区别;其他大模型生成的人格化内容可还原人类心理因子达80%纯度。此外,模型在不同心理压力情境下会调整语言风格,生成的焦虑描述具更低具体性与意象性,符合已有心理学研究结论。结果表明,大模型具备较强但不完全的心理模拟能力,既具替代人类参与研究的潜力,也暴露其认知局限。该工作为量化评估机器心理提供了新路径。
原文摘要 · Abstract (English)
Machine psychology aims to reconstruct the mindset of Large Language Models (LLMs), i.e. how these artificial intelligences perceive and associate ideas. This work introduces PhDGPT, a prompting framework and synthetic dataset that encapsulates the machine psychology of PhD researchers and professors as perceived by OpenAI's GPT-3.5. The dataset consists of 756,000 datapoints, counting 300 iterations repeated across 15 academic events, 2 biological genders, 2 career levels and 42 unique item responses of the Depression, Anxiety, and Stress Scale (DASS-42). PhDGPT integrates these psychometric scores with their explanations in plain language. This synergy of scores and texts offers a dual, comprehensive perspective on the emotional well-being of simulated academics, e.g. male/female PhD students or professors. By combining network psychometrics and psycholinguistic dimensions, this study identifies several similarities and distinctions between human and LLM data. The psychometric networks of simulated male professors do not differ between physical and emotional anxiety subscales, unlike humans. Other LLMs' personification can reconstruct human DASS factors with a purity up to 80%. Furthemore, LLM-generated personifications across different scenarios are found to elicit explanations lower in concreteness and imageability in items coding for anxiety, in agreement with past studies about human psychology. Our findings indicate an advanced yet incomplete ability for LLMs to reproduce the complexity of human psychometric data, unveiling convenient advantages and limitations in using LLMs to replace human participants. PhDGPT also intriguingly capture the ability for LLMs to adapt and change language patterns according to prompted mental distress contextual features, opening new quantitative opportunities for assessing the machine psychology of these artificial intelligences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。