对比GPT-4与GPT-3.5政治倾向,发现前者偏激程度略低且更擅长模仿不同立场。
Is GPT-4 Less Politically Biased than GPT-3.5? A Renewed Investigation of ChatGPT's Political Biases
- 用政治光谱测试和大五人格测试重复100次,量化模型倾向
- GPT-4经济/社会轴得分分别为-5.40/-4.73,比GPT-3.5的-6.59/-6.07更温和
- GPT-4能准确模拟四种政治立场,且测试顺序影响结果体现上下文记忆
本研究调查了ChatGPT的政治偏见与人格特质,重点比较GPT-3.5与GPT-4。通过在每种情景下重复100次执行政治光谱测试(Political Compass Test)和大五人格测试(Big Five Personality Test),获得统计显著结果并分析其相关性。计算均值、标准差并进行显著性检验,以探究两模型间的差异。结果显示,两个模型均呈现进步主义与自由意志主义倾向,但GPT-4的偏见轻微但可忽略地减弱:在经济轴上,GPT-3.5得分为-6.59,而GPT-4为-5.40;社会轴上,分别为-6.07与-4.73。相比之下,GPT-4展现出显著的立场模拟能力,在四类政治定位(自由左、自由右、威权左、威权右)中均准确反映对应象限。在大五人格方面,GPT-3.5表现出高度开放性(85.9%)与宜人性(84.6%),这些特征与人类研究中的自由意志观点相关;而GPT-4整体人格特征较弱,但神经质得分更高。被赋予的政治立场影响了开放性、宜人性和尽责性,再次反映出人类研究中的内在关联性。最后,测试顺序影响模型响应及观察到的相关性,表明存在某种情境记忆。
原文摘要 · Abstract (English)
This work investigates the political biases and personality traits of ChatGPT, specifically comparing GPT-3.5 to GPT-4. In addition, the ability of the models to emulate political viewpoints (e.g., liberal or conservative positions) is analyzed. The Political Compass Test and the Big Five Personality Test were employed 100 times for each scenario, providing statistically significant results and an insight into the results correlations. The responses were analyzed by computing averages, standard deviations, and performing significance tests to investigate differences between GPT-3.5 and GPT-4. Correlations were found for traits that have been shown to be interdependent in human studies. Both models showed a progressive and libertarian political bias, with GPT-4's biases being slightly, but negligibly, less pronounced. Specifically, on the Political Compass, GPT-3.5 scored -6.59 on the economic axis and -6.07 on the social axis, whereas GPT-4 scored -5.40 and -4.73. In contrast to GPT-3.5, GPT-4 showed a remarkable capacity to emulate assigned political viewpoints, accurately reflecting the assigned quadrant (libertarian-left, libertarian-right, authoritarian-left, authoritarian-right) in all four tested instances. On the Big Five Personality Test, GPT-3.5 showed highly pronounced Openness and Agreeableness traits (O: 85.9%, A: 84.6%). Such pronounced traits correlate with libertarian views in human studies. While GPT-4 overall exhibited less pronounced Big Five personality traits, it did show a notably higher Neuroticism score. Assigned political orientations influenced Openness, Agreeableness, and Conscientiousness, again reflecting interdependencies observed in human studies. Finally, we observed that test sequencing affected ChatGPT's responses and the observed correlations, indicating a form of contextual memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。