LLM在决策中模仿人类偏见,但学习方式完全不同。
LLM Agents Display Human Biases but Exhibit Distinct Learning Patterns
- 通过重复选择任务对比人类与LLM行为模式。
- 两者均轻视罕见事件,但成因迥异。
- 人类有近期事件敏感性,而LLM缺乏该机制。
我们研究了大型语言模型(LLMs)在涉及重复选择和从反馈中学习的‘基于经验的决策’任务中的选择模式,并将其与人类参与者的行为进行比较。结果显示,在总体上,LLMs表现出与人类相似的行为偏差:两者均对罕见事件存在低估和相关性效应。然而,对选择模式的更细致分析表明,这种相似性源于截然不同的原因。LLMs表现出强烈的近期偏差,而人类则以更复杂的方式响应。尽管整体行为相似,但对近期事件的依赖程度在两组间差异显著。例如,‘意外触发改变’和‘罕见事件的波浪式近期效应’在人类中稳定存在,但在LLMs中完全缺失。这些发现揭示了使用LLMs模拟和预测人类在学习环境中的行为的局限性,并强调在探究其是否复制人类决策倾向时需进行精细化分析。
原文摘要 · Abstract (English)
We investigate the choice patterns of Large Language Models (LLMs) in the context of Decisions from Experience tasks that involve repeated choice and learning from feedback, and compare their behavior to human participants. We find that on the aggregate, LLMs appear to display behavioral biases similar to humans: both exhibit underweighting rare events and correlation effects. However, more nuanced analyses of the choice patterns reveal that this happens for very different reasons. LLMs exhibit strong recency biases, unlike humans, who appear to respond in more sophisticated ways. While these different processes may lead to similar behavior on average, choice patterns contingent on recent events differ vastly between the two groups. Specifically, phenomena such as ``surprise triggers change" and the ``wavy recency effect of rare events" are robustly observed in humans, but entirely absent in LLMs. Our findings provide insights into the limitations of using LLMs to simulate and predict humans in learning environments and highlight the need for refined analyses of their behavior when investigating whether they replicate human decision making tendencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。