arXiv:2504.20444cs.CLcs.AI2025-04被引 2

测试大模型是否受首因效应影响,发现不同模型反应差异显著。

On Psychology of AI -- Does Primacy Effect Affect ChatGPT and Other LLMs?

  • 用阿希实验范式测试三款大模型对描述顺序的偏好。
  • ChatGPT在同时比较时倾向先正后负的候选人,单独评估则偏好先负后正。
  • 不同模型表现不一,揭示其决策机制差异,适合研究认知偏差的读者。

我们研究了三款商用大模型(ChatGPT、Gemini、Claude)中的首因效应。通过改编阿希(1946)的经典人类实验:给两个候选人提供相同描述,但一个先列正面形容词后列负面,另一个则相反。实验分两种情境:一是将两名候选人同时放入同一提示中,二是分别独立给出。共测试200对候选人。结果显示,在同时比较中,ChatGPT更偏好先正后负的候选人,Gemini则无明显偏好,而Claude拒绝做出选择。在分别评估时,ChatGPT和Claude最常给予同等评价;当有偏好时,两者均更倾向先负后正的候选人,而Gemini则更可能偏好先负后正者。

原文摘要 · Abstract (English)

We study the primacy effect in three commercial LLMs: ChatGPT, Gemini and Claude. We do this by repurposing the famous experiment Asch (1946) conducted using human subjects. The experiment is simple, given two candidates with equal descriptions which one is preferred if one description has positive adjectives first before negative ones and another description has negative adjectives followed by positive ones. We test this in two experiments. In one experiment, LLMs are given both candidates simultaneously in the same prompt, and in another experiment, LLMs are given both candidates separately. We test all the models with 200 candidate pairs. We found that, in the first experiment, ChatGPT preferred the candidate with positive adjectives listed first, while Gemini preferred both equally often. Claude refused to make a choice. In the second experiment, ChatGPT and Claude were most likely to rank both candidates equally. In the case where they did not give an equal rating, both showed a clear preference to a candidate that had negative adjectives listed first. Gemini was most likely to prefer a candidate with negative adjectives listed first.

认知偏差大模型行为首因效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。