arXiv:2412.00323cs.CLcs.AI2024-12综述被引 34

研究大模型认知偏差并测试两种方法的缓解效果

Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments

  • 借鉴众包中的人类缓解方法,测试SoPro与AwaRe在大模型上的适用性
  • GPT-4在应用AwaRe后显著降低六种认知偏差的影响
  • AwaRe有效提升模型决策理性,适合需公平判断的应用场景

大型语言模型(LLMs)基于人类撰写的大量语料训练,表现出优异的任务性能。然而,由于人类易受认知偏差影响,导致非理性判断,LLMs也可能继承此类偏差,引发非理性决策。例如,多项选择题中选项顺序的变化会因顺序偏差影响模型表现。本文首先系统梳理了现有研究中关于LLMs认知偏差及其缓解方法的成果。当前缓解技术存在局限:仅适用于特定偏差,或需长输入输出。随后,我们受众包研究启发,测试了两种针对人类的缓解方法——SoPro与AwaRe在LLMs上的有效性。实验在GPT-3.5和GPT-4上进行,评估六种偏差在应用前后对输出的影响。结果表明,SoPro效果甚微,而AwaRe能有效缓解多种偏差,使模型生成更理性的响应。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are trained on large corpora written by humans and demonstrate high performance on various tasks. However, as humans are susceptible to cognitive biases, which can result in irrational judgments, LLMs can also be influenced by these biases, leading to irrational decision-making. For example, changing the order of options in multiple-choice questions affects the performance of LLMs due to order bias. In our research, we first conducted an extensive survey of existing studies examining LLMs' cognitive biases and their mitigation. The mitigation techniques in LLMs have the disadvantage that they are limited in the type of biases they can apply or require lengthy inputs or outputs. We then examined the effectiveness of two mitigation methods for humans, SoPro and AwaRe, when applied to LLMs, inspired by studies in crowdsourcing. To test the effectiveness of these methods, we conducted experiments on GPT-3.5 and GPT-4 to evaluate the influence of six biases on the outputs before and after applying these methods. The results demonstrate that while SoPro has little effect, AwaRe enables LLMs to mitigate the effect of these biases and make more rational responses.

认知偏差大模型评测缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。