发现大模型在概率推理中同时存在理性与直觉两种思维模式。
Dual Traits in Probabilistic Reasoning of Large Language Models
- 通过实验揭示大模型的后验判断包含符合贝叶斯规则和依赖相似性的双重机制。
- 模型难以回忆基础率信息,且基于提示工程缓解直觉偏差效果有限。
- 研究暗示强化学习中的对比损失可能导致双重判断模式,对高风险应用提出警示。
我们设计了三项实验,探究大语言模型(LLMs)对后验概率的评估方式。结果表明,当前先进模型的后验判断中存在两种并行模式:一种遵循贝叶斯规则的规范性模式,另一种基于相似性的代表性模式,类似于人类的系统1与系统2思维。此外,我们观察到模型难以从记忆中提取基础率信息,且发展提示工程策略以缓解代表性偏差可能极具挑战。我们进一步推测,这种双重判断模式可能是强化学习中人类反馈(RLHF)所采用的对比损失函数所致。研究强调了减少大模型认知偏见的潜在方向,并提醒在关键领域部署时需保持谨慎。
原文摘要 · Abstract (English)
We conducted three experiments to investigate how large language models (LLMs) evaluate posterior probabilities. Our results reveal the coexistence of two modes in posterior judgment among state-of-the-art models: a normative mode, which adheres to Bayes' rule, and a representative-based mode, which relies on similarity -- paralleling human System 1 and System 2 thinking. Additionally, we observed that LLMs struggle to recall base rate information from their memory, and developing prompt engineering strategies to mitigate representative-based judgment may be challenging. We further conjecture that the dual modes of judgment may be a result of the contrastive loss function employed in reinforcement learning from human feedback. Our findings underscore the potential direction for reducing cognitive biases in LLMs and the necessity for cautious deployment of LLMs in critical areas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。